VLDB 2026 Research / reviewers in the wild / expert
Fabio Antonacci
dblp:33/6704
· DBLP profile ↗
64ranked-venue papers
6as first author
23since 2021 · last 2025
0000-0003-4545-0315ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 39 · 4 first-author · 17 since 2021Artificial intelligence and machine learning · 21 · 2 first-author · 5 since 2021Computer networks · 3Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MambaFoley: Foley Sound Generation using Selective State-Space ModelsabstractRecent advancements in deep learning have led to widespread use of techniques for audio content generation, notably employing Denoising Diffusion Probabilistic Models (DDPM) across various tasks. Among these, Foley Sound Synthesis is of particular interest for its role in applications for the creation of multimedia content. Given the temporal-dependent nature of sound, it is crucial to design generative models that can effectively handle the sequential modeling of audio samples. Selective State-Space Models (SSMs) have recently been proposed as a valid alternative to previously proposed techniques, demonstrating competitive performance with lower computational complexity. In this paper, we introduce MambaFoley, a diffusion-based model that, to the best of our knowledge, is the first to leverage the recently proposed SSM known as Mamba for the Foley sound generation task. To evaluate the effectiveness of the proposed method, we compare it with a state-of-the-art Foley sound generative model using both objective and subjective analyses. Marco Furio Colombo, Francesca Ronchini, Luca Comanducci, Fabio Antonacci |
ICASSP | 4 |
| 2025 | A Zero-Shot Physics-Informed Dictionary Learning Approach for Sound Field ReconstructionabstractSound field reconstruction aims to estimate pressure fields in areas lacking direct measurements. Existing techniques often rely on strong assumptions or face challenges related to data availability or the explicit modeling of physical properties. To bridge these gaps, this study introduces a zero-shot, physics-informed dictionary learning approach to perform sound field reconstruction. Our method relies only on a few sparse measurements to learn a dictionary, without the need for additional training data. Moreover, by enforcing the Helmholtz equation during the optimization process, the proposed approach ensures that the reconstructed sound field is represented as a linear combination of a few physically meaningful atoms. Evaluations on real-world data show that our approach achieves comparable performance to state-of-the-art dictionary learning techniques, with the advantage of requiring only a few observations of the sound field and no training on a dataset. Stefano Damiano, Federico Miotello, Mirco Pezzoli, Alberto Bernardini, Fabio Antonacci, Augusto Sarti, Toon van Waterschoot |
ICASSP | 5 |
| 2025 | Past, Present, and Future of Spatial Audio and Room AcousticsabstractThe study of spatial audio and room acoustics aims to create immersive audio experiences by modeling the physics and psychoacoustics of how sound behaves in space. In the long history of this research area, various key technologies have been developed based both on theoretical advancements and practical innovations. We highlight historical achievements, initiative activities, recent advancements, and future outlooks in the research area of spatial audio recording and reproduction, and room acoustic simulation, modeling, analysis, and control. Shoichi Koyama, Enzo De Sena, Prasanga N. Samarasinghe, Mark R. P. Thomas, Fabio Antonacci |
ICASSP | 5 |
| 2025 | Towards HRTF Personalization using Denoising Diffusion ModelsabstractHead-Related Transfer Functions (HRTFs) have fundamental applications for realistic rendering in immersive audio scenarios. However, they are strongly subject-dependent as they vary considerably depending on the shape of the ears, head and torso. Thus, personalization procedures are required for accurate binaural rendering. Recently, Denoising Diffusion Probabilistic Models (DDPMs), a class of generative learning techniques, have been applied to solve a variety of signal processing-related problems. In this paper, we propose a first approach for using DDPM conditioned on anthropometric measurements to generate personalized Head-Related Impulse Response (HRIR), the time-domain representation of HRTF. The results show the feasibility of DDPMs for HRTF personalization obtaining performance in line with state-of-the-art models. Juan Camilo Albarracín Sánchez, Luca Comanducci, Mirco Pezzoli, Fabio Antonacci |
ICASSP | 4 |
| 2024 | Reconstruction of Sound Field Through Diffusion ModelsabstractReconstructing the sound field in a room is an important task for several applications, such as sound control and augmented (AR) or virtual reality (VR). In this paper, we propose a data-driven generative model for reconstructing the magnitude of acoustic fields in rooms with a focus on the modal frequency range. We introduce, for the first time, the use of a conditional Denoising Diffusion Probabilistic Model (DDPM) trained in order to reconstruct the sound field (SF-Diff) over an extended domain. The architecture is devised in order to be conditioned on a set of limited available measurements at different frequencies and generate the sound field in target, unknown, locations. The results show that SF-Diff is able to provide accurate reconstructions. We conduct a comparative analysis with two state-of-the-art baseline methods, one relying on kernel interpolation and the other on deep learning. Federico Miotello, Luca Comanducci, Mirco Pezzoli, Alberto Bernardini, Fabio Antonacci, Augusto Sarti |
ICASSP | 5 |
| 2024 | A Compressive Sensing Approach for the Reconstruction of the Soundfield Produced by Directive Sources in Reverberant RoomsabstractState-of-the-art soundfield reconstruction methods are computationally expensive and their performance is generally undermined by the presence of strong reverberation and of near-field sources, which are usually modeled using omnidirectional radiation patterns. In this work, we propose a compressive sensing approach for the reconstruction of the soundfield produced by arbitrary directive sources in a reverberant room. Assuming sparsity in the distribution of sources, along with a loose prior knowledge on their position and on the geometry of the environment, we reconstruct both the direct and the reverberant components of the soundfield by modeling early reflections as near-field sources. Moreover, the directivity of sources is explicitly modeled by first expressing the soundfield produced by arbitrarily directive sources as an expansion of multipoles, and then introducing group sparsity constraints. Numerical simulations in two rooms with different reverberation characteristics are conducted to perform the validation of the proposed method. Stefano Damiano, Federico Borra, Alberto Bernardini, Fabio Antonacci, Augusto Sarti |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2024 | Deep Prior-Based Audio Inpainting Using Multi-Resolution Harmonic Convolutional Neural NetworksabstractIn this manuscript, we propose a novel method to perform audio inpainting, i.e., the restoration of audio signals presenting multiple missing parts. Audio inpainting can be interpreted in the context of inverse problems as the task of reconstructing an audio signal from its corrupted observation. For this reason, our method is based on a deep prior approach, a recently proposed technique that proved to be effective in the solution of many inverse problems, among which image inpainting. Deep prior allows one to consider the structure of a neural network as an implicit prior and to adopt it as a regularizer. Differently from the classical deep learning paradigm, deep prior performs a single-element training and thus it can be applied to corrupted audio signals independently from the available training data sets. In the context of audio inpainting, a network presenting relevant audio priors will possibly generate a restored version of an audio signal, only provided with its corrupted observation. Our method exploits a time-frequency representation of audio signals and makes use of a multi-resolution convolutional autoencoder, that has been enhanced to perform the harmonic convolution operation. Results show that the proposed technique is able to provide a coherent and meaningful reconstruction of the corrupted audio. It is also able to outperform the methods considered for comparison, in its domain of application. Federico Miotello, Mirco Pezzoli, Luca Comanducci, Fabio Antonacci, Augusto Sarti |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2024 | Acoustic Imaging With Circular Microphone Array: A New Approach for Sound Field AnalysisabstractAcoustic imaging is powerful in collecting spatial information of acoustic sources into a visual representation. In this paper, we focus on the analysis of the exterior acoustic field captured by a circular array of microphones. With a proper parametrization based on angles, we map the directions of arrival of sources as a function of the microphone locations, thus obtaining an acoustic image called “angular space”. Therefore, we introduce a linear transform to enable analysis and synthesis operations for mapping the microphone pressures onto the angular space using local space-time Fourier analysis. We prove the ability of this representation to combine global information coming from multiple arrays in a single acoustic image that can be processed and manipulated. Examples of source localization applications in simulated and measured scenarios show the effectiveness of the proposed method obtaining results comparable with state-of-the-art methods. Marco Olivieri, Amy Bastine, Mirco Pezzoli, Fabio Antonacci, Thushara D. Abhayapala, Augusto Sarti |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2023 | Acoustic Source Localization in the Spherical Harmonics Domain Exploiting Low-Rank ApproximationsabstractAcoustic signal processing in the spherical harmonics domain (SHD) is an active research area that exploits the signals acquired by higher order microphone arrays. A very important task is that concerning the localization of active sound sources. In this paper, we propose a simple yet effective method to localize prominent acoustic sources in adverse acoustic scenarios. By using a proper normalization and arrangement of the estimated spherical harmonic coefficients, we exploit low-rank approximations to estimate the far field modal directional pattern of the dominant source at each time-frame. The experiments confirm the validity of the proposed approach, with superior performance compared to other recent SHD-based approaches. Maximo Cobos, Mirco Pezzoli, Fabio Antonacci, Augusto Sarti |
ICASSP | 3 |
| 2023 | Grad-CAM-Inspired Interpretation of Nearfield Acoustic Holography using Physics-Informed Explainable Neural NetworkabstractThe interpretation and explanation of decision-making processes of neural networks are becoming a key factor in the deep learning field. Although several approaches have been presented for classification problems, the application to regression models needs to be further investigated. In this manuscript we propose a Grad-CAM-inspired approach for the visual explanation of neural network architecture for regression problems. We apply this methodology to a recent physics-informed approach for Nearfield Acoustic Holography, called Kirchhoff-Helmholtz-based Convolutional Neural Network (KHCNN) architecture. We focus on the interpretation of KHCNN using vibrating rectangular plates with different boundary conditions and violin top plates with complex shapes. Results highlight the more informative regions of the input that the network exploits to correctly predict the desired output. The devised approach has been validated in terms of NCC and NMSE using the original input and the filtered one coming from the algorithm. Hagar Kafri, Marco Olivieri, Fabio Antonacci, Mordehay Moradi, Augusto Sarti, Sharon Gannot |
ICASSP | 3 |
| 2023 | Zero-Shot Anomalous Sound Detection in Domestic Environments Using Large-Scale Pretrained Audio Pattern Recognition ModelsabstractAnomalous sound detection is central to audio-based surveillance and monitoring. In a domestic environment, however, the classes of sounds to be considered anomalous are situation-dependent and cannot be determined in advance. At the same time, it is not feasible to expect a demanding labeling effort from the end user. To address these problems, we present a novel zero-shot method relying on an auxiliary large-scale pretrained audio neural network in support of an unsupervised anomaly detector. The auxiliary module is tasked to generate a fingerprint for each sound occasionally registered by the user. These fingerprints are then compared with those extracted from the input audio stream, and the resulting similarity score is used to increase or reduce the sensitivity of the base detector. Experimental results on synthetic data show that the proposed method substantially improves upon the unsupervised base detector and is capable of outperforming existing few-shot learning systems developed for machine condition monitoring without involving additional training. Alessandro Ilic Mezza, Giulio Zanetti, Maximo Cobos, Fabio Antonacci |
ICASSP | 4 |
| 2023 | Real-Time Multichannel Speech Separation and Enhancement Using a Beamspace-Domain-Based Lightweight CNNabstractThe problems of speech separation and enhancement concern the extraction of the speech emitted by a target speaker when placed in a scenario where multiple interfering speakers or noise are present, respectively. A plethora of practical applications such as home assistants and teleconferencing require some sort of speech separation and enhancement pre-processing before applying Automatic Speech Recognition (ASR) systems. In the recent years, most techniques have focused on the application of deep learning to either time-frequency or time-domain representations of the input audio signals. In this paper we propose a real-time multichannel speech separation and enhancement technique, which is based on the combination of a directional representation of the sound field, denoted as beamspace, with a lightweight Convolutional Neural Network (CNN). We consider the case where the Direction-Of-Arrival (DOA) of the target speaker is approximately known, a scenario where the power of the beamspace-based representation can be fully exploited, while we make no assumption regarding the identity of the talker. We present experiments where the model is trained on simulated data and tested on real recordings and we compare the proposed method with a similar state-of-the-art technique. Marco Olivieri, Luca Comanducci, Mirco Pezzoli, Davide Balsarri, Luca Menescardi, Michele Buccoli, Simone Pecorino, Antonio Grosso, Fabio Antonacci, Augusto Sarti |
ICASSP | 9 |
| 2023 | Two-Stage Beamforming With Arbitrary Planar Arrays of Differential Microphone Array UnitsabstractDifferential Microphone Arrays (DMAs) are of great interest in the literature on small-sized microphone arrays, due to their good directivity properties and nearly frequency-invariant spatial responses. Recently developed beamforming techniques combine multiple DMA units to form flexible two-stage spatial filtering systems, where the output of each DMA is fed into a higher-level filter, called virtual filter, for further processing. In this manuscript, we analyze and discuss some properties of a broad class of two-stage beamformers with arbitrary planar geometry. In this context, the DMA units are all assumed to have the same directivity pattern of arbitrary order and can be characterized by a variable number of omnidirectional sensors organized in an arbitrary geometry. For any given choice of the virtual array filter, we introduce a closed-form optimization procedure to design DMA filters that maximize the White Noise Gain (WNG) or the Directivity Factor (DF) of the resulting two-stage beamformer at any frequency. Based on this frequency-dependent design, we propose a frequency-invariant design of the two-stage beamformer and we compare the performance of the two approaches. Finally, we propose two possible computational schemes for the proposed generic two-stage spatial filtering system and discuss their efficiency in performing filtering, steering, and changing beampattern. Davide Albertini, Alberto Bernardini, Federico Borra, Fabio Antonacci, Augusto Sarti |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2022 | On the Prediction of the Frequency Response of a Wooden Plate from Its Mechanical ParametersabstractInspired by deep learning applications in structural mechanics, we focus on how to train two predictors to model the relation between the vibrational response of a prescribed point of a wooden plate and its material properties. In particular, the eigenfrequencies of the plate are estimated via multilinear regression, whereas their amplitude is predicted by a feedforward neural network. We show that labeling the train set by mode numbers instead of by the order of appearance of the eigenfrequencies greatly improves the accuracy of the regression and that the coefficients of the multilinear regressor allow the definition of a linear relation between the first eigenfrequencies of the plate and its material properties. David Giuseppe Badiane, Raffaele Malvermi, Fabio Antonacci, Augusto Sarti |
ICASSP | 4 |
| 2022 | Deepfake Speech Detection Through Emotion Recognition: A Semantic ApproachabstractIn recent years, audio and video deepfake technology has advanced relentlessly, severely impacting people’s reputation and reliability. Several factors have facilitated the growing deepfake threat. On the one hand, the hyper-connected society of social and mass media enables the spread of multimedia content worldwide in real-time, facilitating the dissemination of counterfeit material. On the other hand, neural network-based techniques have made deepfakes easier to produce and difficult to detect, showing that the analysis of low-level features is no longer sufficient for the task. This situation makes it crucial to design systems that allow detecting deepfakes at both video and audio levels. In this paper, we propose a new audio spoofing detection system leveraging emotional features. The rationale behind the proposed method is that audio deepfake techniques cannot correctly synthesize natural emotional behavior. Therefore, we feed our deepfake detector with high-level features obtained from a state-of-the-art Speech Emotion Recognition (SER) system. As the used descriptors capture semantic audio information, the proposed system proves robust in cross-dataset scenarios outperforming the considered baseline on multiple datasets. Emanuele Conti, Davide Salvi, Clara Borrelli, Brian C. Hosler, Paolo Bestagini, Fabio Antonacci, Augusto Sarti, Matthew C. Stamm, Stefano Tubaro |
ICASSP | 6 |
| 2022 | A Data-Driven Approach for Acoustic Parameter Similarity Estimation of Speech RecordingabstractSpeech audio acquisitions exhibit different quality and reverberation properties depending on the recording setup and environment. For this reason, it is expected that speech analysis systems that work correctly on certain audio recordings may fail on others acquired in different acoustic contexts. Therefore, to be able to tell whether a track under analysis shares the same acoustic characteristics of a reference one may be useful to understand if it can be successfully processed by a given speech analysis system. Alternatively, in a forensic scenario, an estimate of acoustic parameter similarity between two tracks can be used to verify whether the recordings have been likely acquired in the same environment or not. In this work, we propose two methods to estimate acoustic parameter similarity between a speech recording under analysis and a reference one. The first method relies on the estimation of channel-based acoustic indicators that are then compared to extract a similarity measure. The second method directly learns a parameter similarity measure through siamese neural networks. Mattia Papa, Clara Borrelli, Paolo Bestagini, Fabio Antonacci, Augusto Sarti, Stefano Tubaro |
ICASSP | 4 |
| 2022 | Sparsity-Based Sound Field Separation in the Spherical Harmonics DomainabstractSound field analysis and reconstruction has been a topic of intense research in the last decades for its multiple applications in spatial audio processing tasks. In this context, the identification of the direct and reverberant sound field components is a problem of great interest, where several solutions exploiting spherical harmonics representations have already been proposed. However, the available techniques demand a large number of high-order microphones (HOMs) and high computational power in order to fulfill the necessary spatial sampling requirements, which can only be reduced by prior information obtained through acoustic measurements. Inspired by compressed sensing approaches, this paper proposes an alternative sparse formulation for estimating the exterior and interior sound field components in the spherical harmonics domain that allows to reduce hardware requirements without the need for additional acoustic measurements. The results show that a considerable reduction in the number of HOMs can be achieved while improving the estimation of the sound field components. Mirco Pezzoli, Maximo Cobos, Fabio Antonacci, Augusto Sarti |
ICASSP | 3 |
| 2022 | Group Dictionary Equivalent Source Method for Sparse Nearfield Acoustic HolographyabstractIn this article we propose a novel methodology for the non-invasive estimation of the geometry of the modes of vibration of a vibrating structure, which uses the signals captured by a microphone array placed in close proximity of the vibrating structure. We propose a measurement approach based on an efficient formulation of the Nearfield Acoustic Holography (NAH) problem, using a small sets of equivalent sources that are able to represent a continuous vibrating structure. In this new reformulation of the solution, a Group Dictionary Equivalent Sources Method is developed, which combines the efficiency and flexibility of equivalent sources, with a sparse reconstruction scheme that takes into account prior knowledge about the mechanical behavior of the object under investigation computed through FEA. Riccardo R. De Lucia, Antonio Canclini, Fabio Antonacci, Augusto Sarti |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2021 | Interpolation of Irregularly Sampled Frequency Response Functions Using Convolutional Neural NetworksabstractIn the field of structural mechanics, classical methods for the vibrational characterization of objects exploit the inherent redundancy of a relevant amount of measurements acquired over regular sampling grids. However, there are cases in which parts of the objects under analysis are not accessible with sensors, leading to irregular sampling grids characterized by holes. Recent works have proved the benefits of adding prior knowledge in these scenarios, either through the definition of a suitable decomposition or using Finite Element modelling. In this paper we propose to use Convolutional Autoencoders (CA) for Frequency Response Function (FRF) interpolation from grids with different subsampling schemes. CA learn a compressed representation from a dataset of FRFs synthetized through Finite Element Analysis. Experiments with numerical and experimental data show the effectiveness of the model with a different amount of missing data and its ability to predict real FRFs characterized by different damping and sampling frequency. Matteo Acerbi, Raffaele Malvermi, Mirco Pezzoli, Fabio Antonacci, Augusto Sarti, Roberto Corradi |
ICASSP | 4 |
| 2021 | Arrays of First-Order Steerable Differential MicrophonesabstractThe literature is rich with techniques for the design of small-size Differential Microphone Arrays (DMAs), known for their almost frequency-invariant beampatterns and low computational cost. Few works, instead, discuss the properties of beamformers based on multiple DMA units. In this paper, we consider arbitrarily shaped planar arrays of DMA units. In turn, each DMA unit is a first-order continuously-steerable differential microphone characterized by an arbitrary configuration of omnidirectional sensors and a symmetric beampattern. We present a beamforming technique that, assumed all the DMA units to steer identical beams in the same direction, allows us to approach the behavior of a Delay-And-Sum beamformer or Super-Directive beamformer by solely varying a single scalar parameter. Efficient implementations of the proposed beamformers can be developed by taking into account that, for a wide range of frequencies, the values of such a parameter are practically invariant with respect to the geometry of the array. Federico Borra, Alberto Bernardini, Ivan Bertuletti, Fabio Antonacci, Augusto Sarti |
ICASSP | 4 |
| 2021 | Sparse Recovery Beamforming and Upscaling in the Ray SpaceabstractWe have been exploring the integration of sparse recovery methods into the ray space transform over the past years and now demonstrate the potential and benefits of beamforming and upscaling signals in the integrated ray space and sparse recovery domain. A primary advantage of the ray space approach derives from its robust ability to integrate information from multiple arrays and viewpoints. Nonetheless, for a given viewpoint, the ray space technique requires a dense array that can be divided into sub-arrays enabling the plenacoustic approach to signal processing. In this work, we explore a method to upscale an array beyond the limits imposed by the inter-microphone distances associated with the array and the concomitant spatial aliasing. In other words, sparse recovery enables one to synthesize or interpolate signals corresponding to an array with a greater number of microphones with a smaller inter-microphone distance. A critical issue is whether or not this interpolative synthesis actually improves array signal processing. This work shows that upscaling signals in the integrated ray space and sparse recovery domain can improve both source localization and separation. Shiduo Yu, Craig T. Jin, Fabio Antonacci, Augusto Sarti |
ICASSP | 3 |
| 2021 | Synthetic speech detection through short-term and long-term prediction tracesabstractAbstract Several methods for synthetic audio speech generation have been developed in the literature through the years. With the great technological advances brought by deep learning, many novel synthetic speech techniques achieving incredible realistic results have been recently proposed. As these methods generate convincing fake human voices, they can be used in a malicious way to negatively impact on today’s society (e.g., people impersonation, fake news spreading, opinion formation). For this reason, the ability of detecting whether a speech recording is synthetic or pristine is becoming an urgent necessity. In this work, we develop a synthetic speech detector. This takes as input an audio recording, extracts a series of hand-crafted features motivated by the speech-processing literature, and classify them in either closed-set or open-set. The proposed detector is validated on a publicly available dataset consisting of 17 synthetic speech generation algorithms ranging from old fashioned vocoders to modern deep learning solutions. Results show that the proposed method outperforms recently proposed detectors in the forensics literature. Clara Borrelli, Paolo Bestagini, Fabio Antonacci, Augusto Sarti, Stefano Tubaro |
EURASIP J. Inf. Secur. | 3 |
| 2021 | Ray-Space-Based Multichannel Nonnegative Matrix Factorization for Audio Source SeparationabstractNonnegative matrix factorization (NMF) has been traditionally considered a promising approach for audio source separation. While standard NMF is only suited for single-channel mixtures, extensions to consider multi-channel data have been also proposed. Among the most popular alternatives, multichannel NMF (MNMF) and further derivations based on constrained spatial covariance models have been successfully employed to separate multi-microphone convolutive mixtures. This letter proposes a MNMF extension by considering a mixture model with Ray-Space-transformed signals, where magnitude data successfully encodes source locations as frequency-independent linear patterns. We show that the MNMF algorithm can be seamlessly adapted to consider Ray-Space-transformed data, providing competitive results with recent state-of-the-art MNMF algorithms in a number of configurations using real recordings. Mirco Pezzoli, Julio J. Carabias-Orti, Maximo Cobos, Fabio Antonacci, Augusto Sarti |
IEEE Signal Process. Lett. | 4 |
| 2020 | Time Difference of Arrival Estimation from Frequency-Sliding Generalized Cross-Correlations Using Convolutional Neural NetworksabstractThe interest in deep learning methods for solving traditional signal processing tasks has been steadily growing in the last years. Time delay estimation (TDE) in adverse scenarios is a challenging problem, where classical approaches based on generalized cross-correlations (GCCs) have been widely used for decades. Recently, the frequency-sliding GCC (FS-GCC) was proposed as a novel technique for TDE based on a sub-band analysis of the cross-power spectrum phase, providing a structured two-dimensional representation of the time delay information contained across different frequency bands. Inspired by deep-learning-based image denoising solutions, we propose in this paper the use of convolutional neural networks (CNNs) to learn the time-delay patterns contained in FS-GCCs extracted in adverse acoustic conditions. Our experiments confirm that the proposed approach provides excellent TDE performance while being able to generalize to different room and sensor setups. Luca Comanducci, Maximo Cobos, Fabio Antonacci, Augusto Sarti |
ICASSP | 3 |
| 2020 | Efficient Implementations of First-Order Steerable Differential Microphone Arrays With Arbitrary Planar GeometryabstractWe present a spatial filtering approach to first-order steerable Differential Microphone Arrays (DMAs) with arbitrary planar geometry. In particular, the design of the spatial filter is based on a recently proposed frequency-domain design methodology that approximates, in a least-square sense, a target beampattern using the Jacobi-Anger expansion involving Bessel functions. Despite the generality of that approach, however, its computational cost turns out to be excessive when working with limited processing resources. The beamforming technique proposed in this manuscript overcomes this issue by exploiting the fact that in DMAs the spacing between sensors is typically smaller than the smallest wavelength of audio signals of interest. This allows us to substitute zero- and first-order Bessel functions with their Taylor series approximation truncated to the first order. Moreover, we show that this approximation allows us to derive an efficient discrete-time-domain implementation of first-order steerable differential beamformers based on arrays with arbitrary geometries. Federico Borra, Alberto Bernardini, Fabio Antonacci, Augusto Sarti |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2020 | A Methodology for the Robust Estimation of the Radiation Pattern of Acoustic SourcesabstractWe propose a novel methodology for estimating the radiation pattern of acoustic sources, which is general enough as to be suitable for a wide variety of sources without the need of anechoic conditions of operation. Multiple plenacoustic cameras (which can be thought of as arrays of acoustic cameras) scan the source while keeping reflections and interferers at bay through deconvolution and windowing of the measured response. In the case of a moving source (e.g. a musical instrument while it is being played), the plenacoustic cameras are also used for tracking the position of the source. As for its orientation, we propose practical solutions for tracking that as well, whenever such information is not known in advance. Two experiments are conducted in order to validate the proposed solution. The former focuses on a commercial loudspeaker cabinet, whose radiation pattern is known in advance and can be used as groundtruth. The latter concerns violins, which exhibit an extremely rich and hard to predict acoustic behavior, due to their inherent structural and constructional complexity. Our method allows us to capture the radiation pattern of the instrument while it is being played, thus returning data corresponding to the natural timbre of the instrument, including the unavoidable acoustic shadow of the violinist's head. Experimental results confirm a relevant improvement in accuracy and robustness afforded by the adoption of dynamic plenacoustic solutions with respect to state-of-the-art techniques. Antonio Canclini, Fabio Antonacci, Stefano Tubaro, Augusto Sarti |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2020 | Frequency-Sliding Generalized Cross-Correlation: A Sub-Band Time Delay Estimation ApproachabstractThe generalized cross-correlation(GCC) is regarded as the most popular approach for estimating the time difference of arrival (TDOA) between the signals received at two sensors. Time delay estimates are obtained by maximizing the GCC output, where the direct-path delay is usually observed as a prominent peak. Moreover, GCCs play also an important role in steered response power (SRP) localization algorithms, where the SRP functional can be written as an accumulation of the GCCs computed from multiple sensor pairs. Unfortunately, the accuracy of TDOA estimates is affected by multiple factors, including noise, reverberation and signal bandwidth. In this paper, a sub-band approach for time delay estimation aimed at improving the performance of the conventional GCC is presented. The proposed method is based on the extraction of multiple GCCs corresponding to different frequency bands of the cross-power spectrum phase in a sliding-window fashion. The major contributions of this paper include: 1) a sub-band GCC representation of the cross-power spectrum phase that, despite having a reduced temporal resolution, provides a more suitable representation for estimating the true TDOA; 2) such matrix representation is shown to be rank one in the ideal noiseless case, a property that is exploited in more adverse scenarios to obtain a more robust and accurate GCC; 3) we propose a set of low-rank approximation alternatives for processing the sub-band GCC matrix, leading to better TDOA estimates and source localization performance. An extensive set of experiments is presented to demonstrate the validity of the proposed approach. Maximo Cobos, Fabio Antonacci, Luca Comanducci, Augusto Sarti |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2020 | Source Localization Using Distributed Microphones in Reverberant Environments Based on Deep Learning and Ray Space TransformabstractIn this article we present a methodology for source localization in reverberant environments from Generalized Cross Correlations (GCCs) computed between spatially distributed individual microphones. Reverberation tends to negatively affect localization based on Time Differences of Arrival (TDOAs), which become inaccurate due to the presence of spurious peaks in the GCC. We therefore adopt a data-driven approach based on a convolutional neural network, which, using the GCCs as input, estimates the source location in two steps. It first computes the Ray Space Transform (RST) from multiple arrays. The RST is a convenient representation of the acoustic rays impinging on the array in a parametric space, called Ray Space. Rays produced by a source are visualized in the RST as patterns, whose position is uniquely related to the source location. The second step consists of estimating the source location through a nonlinear fitting, which estimates the coordinates that best approximate the RST pattern obtained through the first step. It is worth noting that training can be accomplished on simulated data only, thus relaxing the need of actually deploying microphone arrays in the acoustic scene. The localization accuracy of the proposed techniques is similar to the one of SRP-PHAT, however our method demonstrates an increased robustness regarding different distributed microphones configurations. Moreover, the use of the RST as an intermediate representation makes it possible for the network to generalize to data unseen during training. Luca Comanducci, Federico Borra, Paolo Bestagini, Fabio Antonacci, Stefano Tubaro, Augusto Sarti |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2020 | A Parametric Approach to Virtual Miking for Sources of Arbitrary DirectivityabstractIn this article we propose a methodology for the reconstruction of sound fields in arbitrary locations based on the signals acquired by a spatial distribution of compact microphone arrays (virtual miking). The proposed method is suitable for operating in reverberant environments, thanks to a two-stage analysis process, the former of which aims at separating the direct and the diffuse components of the sound field. The method that we propose is inherently parametric, as the sources of the acoustic scene are characterized by parameters describing location and directivity (spherical harmonics expansion), which are extracted from the exterior model of the direct component of the sound field. Once the parameters of the sources are extracted, the direct sound field at an arbitrary location is reconstructed. The diffuse component is reconstructed from the joint knowledge of the diffuse component at the locations of the distributed microphone arrays, under the assumption of isotropic behavior. Results show that the proposed technique is able to analyze the sound field and reconstruct the parameters of the sources that are active in the scene. In addition, the synthesis of the signals at the virtual microphone locations turns out to accurately match (in terms of spatial cues) the actual sound field, as measured by a microphone places in the desired location. Mirco Pezzoli, Federico Borra, Fabio Antonacci, Stefano Tubaro, Augusto Sarti |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2020 | 3D Room Geometry Inference Using a Linear Loudspeaker Array and a Single MicrophoneabstractSound reproduction systems may highly benefit from detailed knowledge of the acoustic space to enhance the spatial sound experience. This article presents a room geometry inference method based on identification of reflective boundaries using a high-resolution direction-of-arrival map produced via room impulse responses (RIRs) measured with a linear loudspeaker array and a single microphone. Exploiting the sparse nature of the early part of the RIRs, Elastic Net regularization is applied to obtain a 2D polar-coordinate map, on which the direct path and early reflections appear as distinct peaks, described by their propagation distance and direction of arrival. Assuming a separable room geometry with four side-walls perpendicular to the floor and ceiling, and imposing pre-defined geometrical constraints on the walls, the 2D-map is segmented into six regions, each corresponding to a particular wall. The salient peaks within each region are selected as candidates for the first-order wall reflections, and a set of potential room geometries is formed by considering all possible combinations of the associated peaks. The room geometry is then inferred using a cost function evaluated on the higher-order reflections computed via beam tracing. The proposed method is tested with both simulated and measured data. Cagdas Tuna, Antonio Canclini, Federico Borra, Philipp Götz, Fabio Antonacci, Andreas Walther 0001, Augusto Sarti, Emanuël A. P. Habets |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2019 | Uniform Linear Arrays of First-Order Steerable Differential MicrophonesabstractWe propose a spatial filtering method for linear arrays of First-Order Steerable Differential Microphones (FOSDMs), which operates in two layers. In the former, signals acquired by individual microphones are locally filtered to produce the outputs of the FOSDMs. In the latter, the outputs of the FOSDMs are processed by another filter. We analyse different design methodologies and study the conditions under which the two filtering layers can be decoupled. The proposed two-layer spatial filter can be flexibly controlled with a single scalar parameter, which can be chosen, for example, to maximize the White Noise Gain (like in a Delay-and-Sum beamformer); or to maximize the Directivity Factor (like in a Super-Directive beamformer); without needing any matrix inversion. The effectiveness of the proposed beamforming method is compared with traditional spatial filtering techniques using different metrics. Federico Borra, Alberto Bernardini, Fabio Antonacci, Augusto Sarti |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2018 | A Weighted Least Squares Beam Shaping Technique for Sound Field ControlabstractA weighted least squares beam shaping technique for sound field control using a loudspeaker array is proposed. Given a desired spatial response at prescribed control points, the space-time filter is designed by solving a least squares minimization problem. To reduce the computational effort, we propose to place control points only along an arc of circumference centered at the center of the array and passing through a region of interest. Furthermore, we adopt a weighted least squares approach for the design of the space-time filter, so that control points at directions towards which we admit a looser control of the sound field are less relevant in the filter design. The choice of the weights depends on the specific application and we demonstrate the feasibility of the proposed approach for sound zones scenario with one bright and one dark zone. Antonio Canclini, Dejan Markovic, Martin Schneider 0009, Fabio Antonacci, Emanuël A. P. Habets, Andreas Walther 0001, Augusto Sarti |
ICASSP | 4 |
| 2018 | Estimation of the Sound Field at Arbitrary Positions in Distributed Microphone Networks Based on Distributed Ray Space TransformabstractIn this paper we propose a parametric sound field reconstruction approach. In particular, the technique is based on the estimation of three parameters for each acoustic source (source position, radiation pattern and source signal) given the signals acquired by few arbitrarily placed microphone arrays. This allows us to synthesize the signal of a virtual microphone placed in any point of the acoustic scene. Mirco Pezzoli, Federico Borra, Fabio Antonacci, Augusto Sarti, Stefano Tubaro |
ICASSP | 3 |
| 2018 | Wave Digital Implementation of Robust First-Order Differential Microphone ArraysabstractIn this letter, a novel time-domain implementation of robust first-order differential microphone arrays (DMAs), based on wave digital filters, is presented. The proposed beamforming method is extremely efficient, as it requires at most two multipliers and one delay for each filter, where the necessary number of filters equals the number of physical microphones of the array, and it avoids the use of fractional delays. The update of the coefficients of the filters, required for reshaping the beampattern, has a significantly lower computational cost with respect to the time-domain methods presented in the literature. This makes the proposed method suitable for real-time DMA applications with time-varying beampatterns. Alberto Bernardini, Fabio Antonacci, Augusto Sarti |
IEEE Signal Process. Lett. | 2 |
| 2017 | Dictionary-based Equivalent Source Method for Near-Field Acoustic HolographyabstractIn this paper, we propose a modification of the standard Equivalent Source Method (ESM) for Near-Field Acoustic Holography (NAH). As in EMS, we aim at modeling the acoustic pressure radiated from a vibrating object, and its surface velocity, as the joint effect of a set of equivalent sources located within or close to the object itself. The estimation of the equivalent source strengths (weigths) comes from the solution of a highly ill-conditioned problem. Rather than solving this problem in the least-squares sense, we exploit the 3D model of the vibrating object, along with a rough estimate of its physical parameters, to restrict the space of the solutions. More specifically, we make use of Finite Element Analysis for populating a compressed dictionary of possible equivalent source weights. NAH is then approached by seeking a sparse linear combination of the entries of the dictionary. Experiments carried on a public database prove the effectiveness of the proposed technique, especially when the number of available microphones is limited, and in the presence of a significant level of measurement noise. Antonio Canclini, Massimo Varini, Fabio Antonacci, Augusto Sarti |
ICASSP | 3 |
| 2017 | Using multi-dimensional correlation for matching and alignment of MoCap and video signalsabstractMotion analysis and tracking often relies on multimodal signals, e.g., video, depth map, motion capture (MoCap), due to the completeness of information they jointly provide. The joint analysis of multimodal signals requires to know the correct timing, i.e., the signals to be aligned. In this paper we propose an approach to automatically estimate the correct matching and alignment between a video and a MoCap recording acquired from the same session, based on the multi-dimensional correlation of velocity-based features extracted from the two recordings. We validate our approach over a dataset of dance recordings of four genres, and we achieve promising results for both the alignment and matching scenarios. Michele Buccoli, Bruno Di Giorgi, Massimiliano Zanoni, Fabio Antonacci, Augusto Sarti |
MMSP | 4 |
| 2017 | Distributed 3D Source Localization from 2D DOA Measurements Using Multiple Linear ArraysabstractThis manuscript addresses the problem of 3D source localization from direction of arrivals (DOAs) in wireless acoustic sensor networks. In this context, multiple sensors measure the DOA of the source, and a central node combines the measurements to yield the source location estimate. Traditional approaches require 3D DOA measurements; that is, each sensor estimates the azimuth and elevation of the source by means of a microphone array, typically in a planar or spherical configuration. The proposed methodology aims at reducing the hardware and computational costs by combining measurements related to 2D DOAs estimated from linear arrays arbitrarily displaced in the 3D space. Each sensor measures the DOA in the plane containing the array and the source. Measurements are then translated into an equivalent planar geometry, in which a set of coplanar equivalent arrays observe the source preserving the original DOAs. This formulation is exploited to define a cost function, whose minimization leads to the source location estimation. An extensive simulation campaign validates the proposed approach and compares its accuracy with state-of-the-art methodologies. Antonio Canclini, Fabio Antonacci, Augusto Sarti, Stefano Tubaro |
Wirel. Commun. Mob. Comput. | 2 |
| 2017 | A Survey of Sound Source Localization Methods in Wireless Acoustic Sensor NetworksabstractWireless acoustic sensor networks (WASNs) are formed by a distributed group of acoustic-sensing devices featuring audio playing and recording capabilities. Current mobile computing platforms offer great possibilities for the design of audio-related applications involving acoustic-sensing nodes. In this context, acoustic source localization is one of the application domains that have attracted the most attention of the research community along the last decades. In general terms, the localization of acoustic sources can be achieved by studying energy and temporal and/or directional features from the incoming sound at different microphones and using a suitable model that relates those features with the spatial location of the source (or sources) of interest. This paper reviews common approaches for source localization in WASNs that are focused on different types of acoustic features, namely, the energy of the incoming signals, their time of arrival (TOA) or time difference of arrival (TDOA), the direction of arrival (DOA), and the steered response power (SRP) resulting from combining multiple microphone signals. Additionally, we discuss methods not only aimed at localizing acoustic sources but also designed to locate the nodes themselves in the network. Finally, we discuss current challenges and frontiers in this field. Maximo Cobos, Fabio Antonacci, Anastasios Alexandridis, Athanasios Mouchtaris, Bowon Lee |
Wirel. Commun. Mob. Comput. | 2 |
| 2017 | Wireless Acoustic Sensor Networks and Applications
Maximo Cobos, Fabio Antonacci, Athanasios Mouchtaris, Bowon Lee |
Wirel. Commun. Mob. Comput. | 2 |
| 2016 | A linear operator for the computation of soundfield mapsabstractIn the process of soundfield imaging, as defined in the literature, a microphone array is subdivided into overlapping sub-arrays and soundfield images are obtained by juxtaposition of spatial spectra computed from individual subarray data. In this paper we show that the whole process can be conveniently seen as a linear transformation applied to array data. This linear transformation embeds a nonlinear mapping to cast the directional information in a more convenient domain: the ray space. We show by simulations that the proposed formulation is suitable for fast implementation of the soundfield imaging operation, and, more specifically, for the localization of acoustic sources. Lucio Bianchi, V. Baldini Anastasio, Dejan Markovic, Fabio Antonacci, Augusto Sarti, Stefano Tubaro |
ICASSP | 4 |
| 2016 | A low-cost solution to 3D pinna modeling for HRTF predictionabstractWe propose an infrared (IR) stereo-vision system for estimating the 3D model of the pinna, based on low-cost devices. A commercial IR calibrated stereo camera is used in conjunction with a structured IR light projector, to acquire highly textured snapshots of the pinna. A point cloud is computed for each snapshot by triangulating the stereo correspondences detected in the acquired IR images. A complete 3D model is computed by aligning and merging the point clouds, and then creating a polygonal mesh surface. The nominal accuracy of the proposed system turns to be about 1 mm, which enables an accurate prediction of the Head Related Transfer Function (HRTF) through numerical acoustic simulation. Luca Bonacina, Antonio Canclini, Fabio Antonacci, Marco Marcon, Augusto Sarti, Stefano Tubaro |
ICASSP | 3 |
| 2016 | Extraction of Acoustic Sources Through the Processing of Sound Field Maps in the Ray SpaceabstractOur goal is to develop a model-based approach to acoustic source extraction from microphone array data, which is suitable for both near-field and far-field sources. A signal representation based on plane-wave (PW) decomposition is suitable for acoustic sources in the far field as the resulting spectrum turns out to be impulsive. When the source approaches the array, however, the curvature of the wavefront causes the spectrum of the PW components to depart from impulsive behavior, thus making source extraction harder to attain. In this paper, we adopt a sound field representation based on the local estimation of the plenacoustic function along the array line. This approach consists of dividing the array into subarrays, and applying the PW analysis on individual subarrays. This has the immediate result of extending the range of validity of the far-field hypothesis, as a source that enters the near-field range of the extended array is still in the far-field range of the subarrays. PW analysis on subarrays allows us to construct the so-called sound field map in a domain of acoustic visibility called ray space. The extraction of the desired source is accomplished through spatial filtering of the sound field map. The design of the spatial filter relies on a linear minimum mean square error criterion defined on the sound field map. The effectiveness of the proposed methodology is proven through an extensive simulation campaign as well as real experiments. Dejan Markovic, Fabio Antonacci, Lucio Bianchi, Stefano Tubaro, Augusto Sarti |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2016 | 3D Beam Tracing Based on Visibility Lookup for Interactive Acoustic ModelingabstractWe present a method for accelerating the computation of specular reflections in complex 3D enclosures, based on acoustic beam tracing. Our method constructs the beam tree on the fly through an iterative lookup process of a precomputed data structure that collects the information on the exact mutual visibility among all reflectors in the environment (region-to-region visibility). This information is encoded in the form of visibility regions that are conveniently represented in the space of acoustic rays using the Plücker coordinates. During the beam tracing phase, the visibility of the environment from the source position (the beam tree) is evaluated by traversing the precomputed visibility data structure and testing the presence of beams inside the visibility regions. The Plücker parameterization simplifies this procedure and reduces its computational burden, as it turns out to be an iterative intersection of linear subspaces. Similarly, during the path determination phase, acoustic paths are found by testing their presence within the nodes of the beam tree data structure. The simulations show that, with an average computation time per beam in the order of a dozen of microseconds, the proposed method can compute a large number of beams at rates suitable for interactive applications with moving sources and receivers. Dejan Markovic, Fabio Antonacci, Augusto Sarti, Stefano Tubaro |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2015 | A Robust and Low-Complexity Source Localization Algorithm for Asynchronous Distributed Microphone NetworksabstractIn this paper, we propose a robust and low-complexity acoustic source localization technique based on time differences of arrival (TDOA), which addresses the scenario of distributed sensor networks in 3D environments. Network nodes are assumed to be unsynchronized, i.e., TDOAs between microphones belonging to different nodes are not available. We begin with showing how to select feasible TDOAs for each sensor node, exploiting both geometrical considerations and a characterization of the overall generalized cross correlation (GCC) shape. We then show how to localize sources in the space-range reference frame, where TDOA measurements have a clear geometrical interpretation that can be fruitfully used in the scenario of unsynchronized sensors. In this framework, in fact, the source corresponds to the apex of a hypercone passing through points described by the sole microphone positions and TDOA measurements. The localization problem is therefore approached as a hypercone fitting problem. Finally, in order to improve the robustness of the estimate, we include an outlier detection procedure based on the evaluation of the hypercone fitting residuals. A refinement of source location estimate is then performed ignoring the contributions coming from outlier measurements. A set of simulations shows the performance of individual blocks of the system, with particular focus on the effect of TDOA selection on source localization and refinement steps. Experiments on real data validate the localization algorithm in an everyday scenario, proving that good accuracy can be obtained while saving computational cost in comparison with state-of-the-art techniques. Antonio Canclini, Paolo Bestagini, Fabio Antonacci, Marco Compagnoni, Augusto Sarti, Stefano Tubaro |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2015 | Multiview Soundfield Imaging in the Projective Ray SpaceabstractA soundfield image is a data structure that efficiently encodes and represents the wave field as captured by a microphone array. Its representation is based on the directional plenacoustic function, which is defined as the radiance of the acoustic paths (rays) that cross the segment that the array lies upon. The soundfield image can be processed “as is” to develop a variety of applications. In its original formulation, the soundfield image is based on a Euclidean parameterization that can accommodate a limited range of rays and is suitable for managing a single array only. In this paper, we generalize this methodology to the case of multiple microphone arrays deployed in space. The use of multiple arrays allows us to capture truly global information on the sound field but requires us to rethink the ray space, and adopt a global representation of the acoustic rays based on projective geometry. After introducing the new parameterization, we present two examples of applications: the estimation of the mutual poses of two or more arrays (self-calibration); and the localization of multiple acoustic sources. The effectiveness of these applications is proven through simulations as well as real data experiments. Dejan Markovic, Fabio Antonacci, Augusto Sarti, Stefano Tubaro |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2014 | Robust beamforming under uncertainties in the loudspeakers directivity patternabstractIn this paper we propose a robust beamforming technique which takes into account uncertainties and variations in the radiation pattern of the loudspeakers. The proposed technique is based on the solution of a robust least-square problem in which the propagation matrix is to some extent unknown. Both simulations and experimental results prove the validity of the proposed methodology in terms of directivity index and white noise gain. Lucio Bianchi, R. Magalotti, Fabio Antonacci, Augusto Sarti, Stefano Tubaro |
ICASSP | 3 |
| 2014 | Estimation of Acoustic Reflection Coefficients Through Pseudospectrum MatchingabstractEstimating the geometric and reflective properties of the environment is important for a wide range of applications of space-time audio processing, from acoustic scene analysis to room equalization and spatial audio rendering. In this manuscript, we propose a methodology for frequency-subband in-situ estimation of the reflection coefficients of planar surfaces. This is a rather challenging task, as the reflection coefficients depend on the frequency and the angle of incidence and their estimate is highly sensitive to background noise and interfering sources. Our method is based on the assumption that we know the geometry of the reflectors; the position and the radiation pattern of the source; the position and the spatial response of the array. Applying beamforming algorithms on a single set of measured sensor data, we estimate the angular distribution of the acoustic energy (angular pseudospectrum) that impinges on a microphone array. We then apply a two-step iterative estimation technique based on an Expectation-Maximization (EM) algorithm. The first step estimates the scaling factors. The second one infers the reflection coefficients from the scaling factors. Under the assumption of additive white Gaussian noise, we finally determine the reflection coefficients with a Maximum Likelihood (ML) estimation method. The effectiveness and the accuracy of the proposed technique are assessed through experiments based on measured data. Dejan Markovic, Konrad Kowalczyk, Fabio Antonacci, Christian Hofmann 0001, Augusto Sarti, Walter Kellermann |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2013 | Localization of virtual acoustic sources based on the Hough transform for sound field rendering applicationsabstractIn this paper we propose a methodology for the localization of virtual acoustic sources for sound field rendering applications. After the reconstruction of the sound field in the listening area by means of circular harmonic decomposition, the virtual source location is found through the Hough transform. We prove the accuracy of the proposed methodology by comparing the source locations estimates with those of a subjective test campaign. Lucio Bianchi, Fabio Antonacci, Antonio Canclini, Augusto Sarti, Stefano Tubaro |
ICASSP | 2 |
| 2013 | Rendering of directional sources through loudspeaker arrays based on plane wave decompositionabstractIn this paper we present a technique for the rendering of directional sources by means of loudspeaker arrays. The proposed methodology is based on a decomposition of the sound field in terms of plane waves. Within this framework the directivity of the source is naturally included in the rendering problem, therefore accommodating the directivity into the picture becomes much simpler. For this purpose, the loudspeaker array is subdivided into overlapping sub-arrays, each generating a plane wave component. The individual plane waves are then weighed by the desired directivity pattern. Simulations and experimental results show that the proposed technique is able to reproduce the sound field of directional sources with an improved accuracy with respect to existing techniques. Lucio Bianchi, Fabio Antonacci, Augusto Sarti, Stefano Tubaro |
MMSP | 2 |
| 2013 | Acoustic Source Localization With Distributed Asynchronous Microphone NetworksabstractWe propose a method for localizing an acoustic source with distributed microphone networks. Time Differences of Arrival (TDOAs) of signals pertaining the same sensor are estimated through Generalized Cross-Correlation. After a TDOA filtering stage that discards measurements that are potentially unreliable, source localization is performed by minimizing a fourth-order polynomial that combines hyperbolic constraints from multiple sensors. The algorithm turns to exhibit a significantly lower computational cost compared with state-of-the-art techniques, while retaining an excellent localization accuracy in fairly reverberant conditions. Antonio Canclini, Fabio Antonacci, Augusto Sarti, Stefano Tubaro |
IEEE Trans. Speech Audio Process. | 2 |
| 2013 | Soundfield Imaging in the Ray SpaceabstractIn this work we propose a general approach to acoustic scene analysis based on a novel data structure (ray-space image) that encodes the directional plenacoustic function over a line segment (Observation Window, OW). We define and describe a system for acquiring a ray-space image using a microphone array and refer to it as ray-space (or “soundfield”) camera. The method consists of acquiring the pseudo-spectra corresponding to a grid of sampling points over the OW, and remapping them onto the ray space, which parameterizes acoustic paths crossing the OW. The resulting ray-space image displays the information gathered by the sensors in such a way that the elements of the acoustic scene (sources and reflectors) will be easy to discern, recognize and extract. The key advantage of this method is that ray-space images, irrespective of the application, are generated by a common (and highly parallelizable) processing layer, and can be processed using methods coming from the extensive literature of pattern analysis. After defining the ideal ray-space image in terms of the directional plenacoustic function, we show how to acquire it using a microphone array. We also discuss resolution and aliasing issues and show two simple examples of applications of ray-space imaging. Dejan Markovic, Fabio Antonacci, Augusto Sarti, Stefano Tubaro |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2012 | Inference of Room Geometry From Acoustic Impulse ResponsesabstractAcoustic scene reconstruction is a process that aims to infer characteristics of the environment from acoustic measurements. We investigate the problem of locating planar reflectors in rooms, such as walls and furniture, from signals obtained using distributed microphones. Specifically, localization of multiple two- dimensional (2-D) reflectors is achieved by estimation of the time of arrival (TOA) of reflected signals by analysis of acoustic impulse responses (AIRs). The estimated TOAs are converted into elliptical constraints about the location of the line reflector, which is then localized by combining multiple constraints. When multiple walls are present in the acoustic scene, an ambiguity problem arises, which we show can be addressed using the Hough transform. Additionally, the Hough transform significantly improves the robustness of the estimation for noisy measurements. The proposed approach is evaluated using simulated rooms under a variety of different controlled conditions where the floor and ceiling are perfectly absorbing. Results using AIRs measured in a real environment are also given. Additionally, results showing the robustness to additive noise in the TOA information are presented, with particular reference to the improvement achieved through the use of the Hough transform. Fabio Antonacci, Jason Filos, Mark R. P. Thomas, Emanuël A. P. Habets, Augusto Sarti, Patrick A. Naylor, Stefano Tubaro |
IEEE Trans. Speech Audio Process. | 1 |
| 2012 | Localization of Acoustic Sources Through the Fitting of Propagation Cones Using Multiple Independent ArraysabstractIn this paper, we propose a novel acoustic source localization method that accommodates the general scenario of multiple independent microphone arrays. The method is based on a 3-D parameter space defined by the 2-D spatial location of a source and the range difference extracted from the time difference of arrival (TDOA). In this space, the set of points that correspond to a given range lie on a circle that expand as the range increases, forming a cone whose apex is the actual location of the source. In this parameter space, the lack of synchronization between arrays results in the fact that clusters of data associated to individual arrays are free to shift along the range axis. The cone constraint, in fact, enables the realignment of such clusters while positioning the cone vertex (source location), thus resulting in a joint data re-synchronization and source localization. We also propose a novel and general analysis methodology for swiftly assessing the localization error as a function of the TDOA uncertainties, which is remarkably accurate for small localization bias. With the aid of this method, simulations and experiments on real data, we show that the cone-fitting process offers excellent localization accuracy in the scenario of multiple unsynchronized arrays, as well as in simpler single-array scenarios, also in comparison with state-of-the-art techniques. We also show that the proposed method offers the desired flexibility for adapting to arbitrary geometries of microphone clusters. Marco Compagnoni, Paolo Bestagini, Fabio Antonacci, Augusto Sarti, Stefano Tubaro |
IEEE Trans. Speech Audio Process. | 3 |
| 2011 | A methodology for evaluating the accuracy of wave field rendering techniquesabstractIn this paper we propose a methodology for assessing the accuracy of techniques of wave field rendering through loudspeaker arrays. In order to measure the rendered wave field we adopt a solution based on a circular harmonic analysis of the sound field captured by a virtual microphone array. As a result of this analysis stage, we are able to compare the target, the theoretical and the measured wave fields, which may differ due to the non-ideality in the loudspeaker array or in the environment that generates some spurious reverberations. Moreover, in order to quantify the error between target, theoretical and measured wave fields, we define some evaluation metrics, based on RMSE and modal analysis of the acquired wave fields. We show some experimental results on real data. Antonio Canclini, Paolo Annibale, Fabio Antonacci, Augusto Sarti, Rudolf Rabenstein, Stefano Tubaro |
ICASSP | 3 |
| 2011 | From direction of arrival estimates to localization of planar reflectors in a two dimensional geometryabstractIn this paper we propose a novel technique to localize planar obstacles through the measurement of the Direction of Arrival by a microphone array. The measurement of the Direction of Arrival of the reflected path is turned into a quadratic constraint where the unknowns are the line parameters of the reflector. A cost function that combines multiple constraints is then derived. A parametric description of the obstacle is found by minimization of the cost function. Some simulations and experimental results show the feasibility of the proposed method. Antonio Canclini, Paolo Annibale, Fabio Antonacci, Augusto Sarti, Rudolf Rabenstein, Stefano Tubaro |
ICASSP | 3 |
| 2010 | Geometric reconstruction of the environment from its response to multiple acoustic emissionsabstractIn this paper we propose a method for reconstructing the 2D geometry of the surrounding environment based on the signals acquired by a fixed microphone, when a series of acoustic stimula are produced in different positions in space. After estimating the Times Of Arrival (TOAs) of the reflective paths, we turn each TOA into a projective geometric constraint that can be used for determining the locations of the reflectors. The result consists of a collection of planar surfaces that correspond to the reflectors' locations. In this paper we present the whole processing chain and prove its effectiveness through experimental results. Fabio Antonacci, Augusto Sarti, Stefano Tubaro |
ICASSP | 1 |
| 2010 | Visibility-based beam tracing for soundfield renderingabstractIn this paper we present a visibility-based beam tracing solution for the simulation of the acoustics of environment that makes use of a projective geometry representation. More specifically, projective geometry turns out to be useful for the pre-computation of the visibility among all the reflectors in the environment. The simulation engine has a straightforward application in the rendering of the acoustics of virtual environments using loudspeaker arrays. More specifically, the acoustic wavefield is conceived as a superposition of acoustic beams, whose parameters (i.e. origin, orientation and aperture) are computed using the fast beam tracing methodology presented here. This information is processed by the rendering engine to compute spatial filters to be applied to the loudspeakers within the array. Simulative results show that an accurate simulation of the acoustic wavefield can be obtained using this approach. Dejan Markovic, Antonio Canclini, Fabio Antonacci, Augusto Sarti, Stefano Tubaro |
MMSP | 3 |
| 2010 | Geometric calibration of distributed microphone arrays from acoustic source correspondencesabstractThis paper proposes a method that solves the problem of geometric calibration of microphone arrays. We consider a distributed system, in which each array is controlled by separate acquisition devices that do not share a common synchronization clock. Given a set of probing sources, e.g. loudspeakers, each array computes an estimate of the source locations using a conventional TDOA-based algorithm. These observations are fused together by the proposed method, in order to estimate the position and pose of one array with respect to the other. Unlike previous approaches, we explicitly consider the anisotropic distribution of localization errors. As such, the proposed method is able to address the problem of geometric calibration when the probing sources are located both in the near- and far-field of the microphone arrays. Experimental results demonstrate that the improvement in terms of calibration accuracy with respect to state-of-the-art algorithms can be substantial, especially in the far-field. S. Daniele Valente, Marco Tagliasacchi, Fabio Antonacci, Paolo Bestagini, Augusto Sarti, Stefano Tubaro |
MMSP | 3 |
| 2009 | Geometric calibration of distributed microphone arraysabstractComputational auditory scene analysis exploits signals acquired by means of microphone arrays. In some circumstances, more than one array is deployed in the same environment. In order to effectively fuse the information gathered by each array, the relative location and pose of the arrays needs to be obtained solving a problem of geometric inter-array calibration. We consider the case where the arrays do not share a synchronous clock, which impairs the use of time-difference of arrival measures across arrays. Conversely, each array produces an acoustic image, which describes the energy of acoustic signals received from different directions. We jointly consider acoustic images acquired by the different arrays and adapt computer vision techniques to solve the calibration problem, thus estimating the location and pose of microphone arrays sensing the same auditory scene. We evaluate the robustness of the calibration process in a simulated environment and we investigate the effect of the various system parameters, namely the number of probing signal locations, the resolution of the acoustic images, the non-ideal intra-array calibration. Alessandro Redondi, Marco Tagliasacchi, Fabio Antonacci, Augusto Sarti |
MMSP | 3 |
| 2008 | Fast Tracing of Acoustic Beams and Paths Through Visibility LookupabstractThe beam tracing method can be used for the fast tracing of a large number of acoustic paths through a direct lookup of a special tree-like data structure (beam tree) that describes the iterated visibility information from one specific position. This structure describes the branching of bundles of rays (beams) as they encounter reflectors in their paths. For this reason, beam tracing is suitable for real-time acoustic rendering even when the receiver is moving. In this paper, we propose a novel technique that enables the fast tracing of a large number of acoustic beams through the iterative lookup of a special data structure that describes the global visibility between reflectors. The method enables the immediate generation of the beam tree corresponding to an arbitrary source location, which can then be used for path tracing through direct lookup. In practice, this technique generalizes the traditional beam-tracing method as it makes it suitable for real-time acoustic rendering not just when the receiver is moving but also when the source is moving. The method enables real-time modeling of acoustic propagation and real-time auralization in complex 2-D and 2-Dtimes1-D environments (e.g., vertical walls limited by horizontal floor and ceiling), which makes it suitable for applications of real-time virtual acoustics, immersive gaming, and advanced acoustic rendering. Some experimental results show the effectiveness of fast beam tracing with respect to the state of the art in acoustic beam tracing. Fabio Antonacci, Marco Foco, Augusto Sarti, Stefano Tubaro |
IEEE Trans. Speech Audio Process. | 1 |
| 2007 | Tracking of two acoustic sources in reverberant environments using a particle swarm optimizerabstractIn this paper we consider the problem of tracking multiple acoustic sources in reverberant environments. The solution that we propose is based on the combination of two techniques. A blind source separation (BSS) method known as TRINICON [5] is applied to the signals acquired by the microphone arrays. The TRINICON de-mixing filters are used to obtain the Time Differences of Arrival (TDOAs), which are related to the source location through a nonlinear function. A particle filter is then applied in order to localize the sources. Particles move according to a swarm-like dynamics, which significatively reduces the number of particles involved with respect to traditional particle filter. We discuss results for the case of two sources and four microphone pairs. In addition, we propose a method, based on detecting source inactivity, which overcomes the ambiguities that intrinsically arise when only two microphone pairs are used. Experimental results demonstrate that the average localization error on a variety of pseudo-random trajectories is around 40 cm when the T60reverberation time is 0.6s. Fabio Antonacci, Davide Riva, Augusto Sarti, Marco Tagliasacchi, Stefano Tubaro |
AVSS | 1 |
| 2007 | Scream and gunshot detection and localization for audio-surveillance systemsabstractThis paper describes an audio-based video surveillance system which automatically detects anomalous audio events in a public square, such as screams or gunshots, and localizes the position of the acoustic source, in such a way that a video-camera is steered consequently. The system employs two parallel GMM classifiers for discriminating screams from noise and gunshots from noise, respectively. Each classifier is trained using different features, chosen from a set of both conventional and innovative audio features. The location of the acoustic source which has produced the sound event is estimated by computing the time difference of arrivals of the signal at a microphone array and using linear-correction least square localization algorithm. Experimental results show that our system can detect events with a precision of 93% at a false rejection rate of 5% when the SNR is 10dB, while the source direction can be estimated with a precision of one degree. A real-time implementation of the system is going to be installed in a public square of Milan. Giuseppe Valenzise, Luigi Gerosa, Marco Tagliasacchi, Fabio Antonacci, Augusto Sarti |
AVSS | 4 |
| 2005 | Efficient source localization and tracking in reverberant environments using microphone arraysabstractIn this paper, we propose an algorithm for acoustic source localization and tracking that is suitable for reverberant environments. The approach that we propose is based on the iterative identification of the FIR channels that link source and microphones through an LMS method (multi-channel LMS), but we propose additional solutions that significantly improve this method in terms of computational efficiency and localization reliability, without affecting its convergence properties. This is achieved through a modified block-wise implementation of the approach combined with Kalman filtering. We also show the results of extensive comparative testing using novel performance parameters for the assessment of localization reliability. Fabio Antonacci, Davide Lonoce, Marco Motta, Augusto Sarti, Stefano Tubaro |
ICASSP (4) | 1 |
| 2004 | Accurate and fast audio-realistic rendering of sounds in virtual environmentsabstractIn this paper we propose a novel method for real-time auralization of sounds in complex environments using visibility diagrams. The method accounts for both specular and diffracted reflections with receivers and sources that are free to move. Our solution concentrates in a pre-processing phase all the operations that can be conducted without knowledge of either source or receiver locations. In fact we pre-compute a set of visibility diagrams and diffracted beam trees. Once the source location is specified, we can determine the reflective beam trees through a simple lookup process on the visibility diagrams. The additional knowledge of the receiver location allows us to immediately generate all reflective and diffractive paths that link source and receiver. Fabio Antonacci, Marco Foco, Augusto Sarti, Stefano Tubaro |
MMSP | 1 |