EDBT 2026 Demo / reviewers in the wild / expert
Augusto Sarti
dblp:11/1386
· DBLP profile ↗
131ranked-venue papers
5as first author
31since 2021 · last 2025
0000-0002-5803-1702ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 93 · 5 first-author · 21 since 2021Artificial intelligence and machine learning · 32 · 8 since 2021Computer networks · 3Systems, architecture and hardware · 2 · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Zero-Shot Physics-Informed Dictionary Learning Approach for Sound Field ReconstructionabstractSound field reconstruction aims to estimate pressure fields in areas lacking direct measurements. Existing techniques often rely on strong assumptions or face challenges related to data availability or the explicit modeling of physical properties. To bridge these gaps, this study introduces a zero-shot, physics-informed dictionary learning approach to perform sound field reconstruction. Our method relies only on a few sparse measurements to learn a dictionary, without the need for additional training data. Moreover, by enforcing the Helmholtz equation during the optimization process, the proposed approach ensures that the reconstructed sound field is represented as a linear combination of a few physically meaningful atoms. Evaluations on real-world data show that our approach achieves comparable performance to state-of-the-art dictionary learning techniques, with the advantage of requiring only a few observations of the sound field and no training on a dataset. Stefano Damiano, Federico Miotello, Mirco Pezzoli, Alberto Bernardini, Fabio Antonacci, Augusto Sarti, Toon van Waterschoot |
ICASSP | 6 |
| 2024 | Reconstruction of Sound Field Through Diffusion ModelsabstractReconstructing the sound field in a room is an important task for several applications, such as sound control and augmented (AR) or virtual reality (VR). In this paper, we propose a data-driven generative model for reconstructing the magnitude of acoustic fields in rooms with a focus on the modal frequency range. We introduce, for the first time, the use of a conditional Denoising Diffusion Probabilistic Model (DDPM) trained in order to reconstruct the sound field (SF-Diff) over an extended domain. The architecture is devised in order to be conditioned on a set of limited available measurements at different frequencies and generate the sound field in target, unknown, locations. The results show that SF-Diff is able to provide accurate reconstructions. We conduct a comparative analysis with two state-of-the-art baseline methods, one relying on kernel interpolation and the other on deep learning. Federico Miotello, Luca Comanducci, Mirco Pezzoli, Alberto Bernardini, Fabio Antonacci, Augusto Sarti |
ICASSP | 6 |
| 2024 | Toward deep drum source separationabstractIn the past, the field of drum source separation faced significant challenges due to limited data availability, hindering the adoption of cutting-edge deep learning methods that have found success in other related audio applications. In this letter, we introduce StemGMD, a large- scale audio dataset of isolated single-instrument drum stems. Each audio clip is synthesized from MIDI recordings of expressive drum performances using ten real-sounding acoustic drum kits. Totaling 1224 h, StemGMD is the largest audio dataset of drums to date and the first to comprise isolated audio clips for every instrument in a canonical nine-piece drum kit. We leverage StemGMD to develop LarsNet, a novel deep drum source separation model. Through a bank of dedicated U-Nets, LarsNet can separate five stems from a stereo drum mixture faster than real-time and is shown to considerably outperform state-of-the-art nonnegative spectro-temporal factorization methods. Alessandro Ilic Mezza, Riccardo Giampiccolo, Alberto Bernardini, Augusto Sarti |
Pattern Recognit. Lett. | 4 |
| 2024 | A Compressive Sensing Approach for the Reconstruction of the Soundfield Produced by Directive Sources in Reverberant RoomsabstractState-of-the-art soundfield reconstruction methods are computationally expensive and their performance is generally undermined by the presence of strong reverberation and of near-field sources, which are usually modeled using omnidirectional radiation patterns. In this work, we propose a compressive sensing approach for the reconstruction of the soundfield produced by arbitrary directive sources in a reverberant room. Assuming sparsity in the distribution of sources, along with a loose prior knowledge on their position and on the geometry of the environment, we reconstruct both the direct and the reverberant components of the soundfield by modeling early reflections as near-field sources. Moreover, the directivity of sources is explicitly modeled by first expressing the soundfield produced by arbitrarily directive sources as an expansion of multipoles, and then introducing group sparsity constraints. Numerical simulations in two rooms with different reverberation characteristics are conducted to perform the validation of the proposed method. Stefano Damiano, Federico Borra, Alberto Bernardini, Fabio Antonacci, Augusto Sarti |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2024 | Deep Prior-Based Audio Inpainting Using Multi-Resolution Harmonic Convolutional Neural NetworksabstractIn this manuscript, we propose a novel method to perform audio inpainting, i.e., the restoration of audio signals presenting multiple missing parts. Audio inpainting can be interpreted in the context of inverse problems as the task of reconstructing an audio signal from its corrupted observation. For this reason, our method is based on a deep prior approach, a recently proposed technique that proved to be effective in the solution of many inverse problems, among which image inpainting. Deep prior allows one to consider the structure of a neural network as an implicit prior and to adopt it as a regularizer. Differently from the classical deep learning paradigm, deep prior performs a single-element training and thus it can be applied to corrupted audio signals independently from the available training data sets. In the context of audio inpainting, a network presenting relevant audio priors will possibly generate a restored version of an audio signal, only provided with its corrupted observation. Our method exploits a time-frequency representation of audio signals and makes use of a multi-resolution convolutional autoencoder, that has been enhanced to perform the harmonic convolution operation. Results show that the proposed technique is able to provide a coherent and meaningful reconstruction of the corrupted audio. It is also able to outperform the methods considered for comparison, in its domain of application. Federico Miotello, Mirco Pezzoli, Luca Comanducci, Fabio Antonacci, Augusto Sarti |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2024 | Acoustic Imaging With Circular Microphone Array: A New Approach for Sound Field AnalysisabstractAcoustic imaging is powerful in collecting spatial information of acoustic sources into a visual representation. In this paper, we focus on the analysis of the exterior acoustic field captured by a circular array of microphones. With a proper parametrization based on angles, we map the directions of arrival of sources as a function of the microphone locations, thus obtaining an acoustic image called “angular space”. Therefore, we introduce a linear transform to enable analysis and synthesis operations for mapping the microphone pressures onto the angular space using local space-time Fourier analysis. We prove the ability of this representation to combine global information coming from multiple arrays in a single acoustic image that can be processed and manipulated. Examples of source localization applications in simulated and measured scenarios show the effectiveness of the proposed method obtaining results comparable with state-of-the-art methods. Marco Olivieri, Amy Bastine, Mirco Pezzoli, Fabio Antonacci, Thushara D. Abhayapala, Augusto Sarti |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2023 | Diffusion-Based Sound Source Localization Using Networks of Planar Microphone ArraysabstractIn this work, we propose a novel approach for distributed 3D sound source localization and tracking based on networks of planar microphone arrays, each of which estimates a 2D Direction Of Arrival (DOA). The proposed method is computationally distributed and eliminates the need for a specialized node to collect and process all information. Sound source localization is achieved by considering the task as a distributed optimization problem approached using the Adapt Then Combine (ATC) diffusion technique. This approach also allows the development of cooperation strategies between sensor nodes (i.e., microphone arrays). We propose the use of a cooperation strategy that improves the localization accuracy by exploiting the estimated error statistics of each sensor node and penalizing the noisy arrays. We then evaluate the proposed approach in terms of localization accuracy and robustness to noisy sensor measurements. Davide Albertini, Gioele Greco, Alberto Bernardini, Augusto Sarti |
ICASSP | 4 |
| 2023 | Acoustic Source Localization in the Spherical Harmonics Domain Exploiting Low-Rank ApproximationsabstractAcoustic signal processing in the spherical harmonics domain (SHD) is an active research area that exploits the signals acquired by higher order microphone arrays. A very important task is that concerning the localization of active sound sources. In this paper, we propose a simple yet effective method to localize prominent acoustic sources in adverse acoustic scenarios. By using a proper normalization and arrangement of the estimated spherical harmonic coefficients, we exploit low-rank approximations to estimate the far field modal directional pattern of the dominant source at each time-frame. The experiments confirm the validity of the proposed approach, with superior performance compared to other recent SHD-based approaches. Maximo Cobos, Mirco Pezzoli, Fabio Antonacci, Augusto Sarti |
ICASSP | 4 |
| 2023 | Grad-CAM-Inspired Interpretation of Nearfield Acoustic Holography using Physics-Informed Explainable Neural NetworkabstractThe interpretation and explanation of decision-making processes of neural networks are becoming a key factor in the deep learning field. Although several approaches have been presented for classification problems, the application to regression models needs to be further investigated. In this manuscript we propose a Grad-CAM-inspired approach for the visual explanation of neural network architecture for regression problems. We apply this methodology to a recent physics-informed approach for Nearfield Acoustic Holography, called Kirchhoff-Helmholtz-based Convolutional Neural Network (KHCNN) architecture. We focus on the interpretation of KHCNN using vibrating rectangular plates with different boundary conditions and violin top plates with complex shapes. Results highlight the more informative regions of the input that the network exploits to correctly predict the desired output. The devised approach has been validated in terms of NCC and NMSE using the original input and the filtered one coming from the algorithm. Hagar Kafri, Marco Olivieri, Fabio Antonacci, Mordehay Moradi, Augusto Sarti, Sharon Gannot |
ICASSP | 5 |
| 2023 | Real-Time Multichannel Speech Separation and Enhancement Using a Beamspace-Domain-Based Lightweight CNNabstractThe problems of speech separation and enhancement concern the extraction of the speech emitted by a target speaker when placed in a scenario where multiple interfering speakers or noise are present, respectively. A plethora of practical applications such as home assistants and teleconferencing require some sort of speech separation and enhancement pre-processing before applying Automatic Speech Recognition (ASR) systems. In the recent years, most techniques have focused on the application of deep learning to either time-frequency or time-domain representations of the input audio signals. In this paper we propose a real-time multichannel speech separation and enhancement technique, which is based on the combination of a directional representation of the sound field, denoted as beamspace, with a lightweight Convolutional Neural Network (CNN). We consider the case where the Direction-Of-Arrival (DOA) of the target speaker is approximately known, a scenario where the power of the beamspace-based representation can be fully exploited, while we make no assumption regarding the identity of the talker. We present experiments where the model is trained on simulated data and tested on real recordings and we compare the proposed method with a similar state-of-the-art technique. Marco Olivieri, Luca Comanducci, Mirco Pezzoli, Davide Balsarri, Luca Menescardi, Michele Buccoli, Simone Pecorino, Antonio Grosso, Fabio Antonacci, Augusto Sarti |
ICASSP | 10 |
| 2023 | Towards a general framework for the annotation of dance motion sequences
Katerina El Raheb, Michele Buccoli, Massimiliano Zanoni, Akrivi Katifori, Aristotelis Kasomoulis, Augusto Sarti, Yannis E. Ioannidis |
Multim. Tools Appl. | 6 |
| 2023 | Loudspeaker virtualization-Part II: The inverse transducer model and the Direct-Inverse-Direct Chain
Alberto Bernardini, Lucio Bianchi, Augusto Sarti |
Signal Process. | 3 |
| 2023 | Loudspeaker virtualization-Part I: Digital modeling and implementation of the nonlinear transducer equivalent circuit
Alberto Bernardini, Lucio Bianchi, Augusto Sarti |
Signal Process. | 3 |
| 2023 | Virtualization of Guitar Pickups Through Circuit InversionabstractA method for circuit system inversion has been recently employed to develop digital algorithms for loudspeaker virtualization. In this brief, building on such promising results, we propose an algorithm that flips the paradigm by virtualizing sensors rather than actuators. In particular, we derive direct and inverse nonlinear circuital models of guitar pickup systems by assuming string vertical excitations, and we then present a virtualization algorithm based on circuit inversion. The proposed circuital models are then implemented in the discrete-time domain in a fully explicit fashion (i.e., with no use of iterative solvers) by employing Wave Digital Filter principles. Finally, we validate and test the designed algorithm on the physical output voltage of a guitar pickup system, making it sound as if acquired by two other magnetic pickups characterized by a different nonlinear behavior. Riccardo Giampiccolo, Alberto Bernardini, Augusto Sarti |
IEEE Signal Process. Lett. | 3 |
| 2023 | Virtual Bass Enhancement via Music DemixingabstractVirtual Bass Enhancement (VBE) refers to a class of digital signal processing algorithms that aim at enhancing the perception of low frequencies in audio applications. Such algorithms typically exploit well-known psychoacoustic effects and are particularly valuable for improving the performance of small-size transducers often found in consumer electronics. Though both time- and frequency-domain techniques have been proposed in the literature, none of them capitalizes on the latest achievements of deep learning as far as music processing is concerned. In this letter, we propose a novel time-domain VBE algorithm that incorporates a deep neural network for music demixing as part of the processing pipeline. This technique is shown to improve the bass perception and reduce inharmonic distortion, i.e., the main issue of existing time-domain VBE algorithms. The results of a perceptual test are then presented, showing that the proposed method is able to outperform state-of-the-art algorithms both in terms of bass enhancement and basic audio quality. Riccardo Giampiccolo, Alessandro Ilic Mezza, Alberto Bernardini, Augusto Sarti |
IEEE Signal Process. Lett. | 4 |
| 2023 | Two-Stage Beamforming With Arbitrary Planar Arrays of Differential Microphone Array UnitsabstractDifferential Microphone Arrays (DMAs) are of great interest in the literature on small-sized microphone arrays, due to their good directivity properties and nearly frequency-invariant spatial responses. Recently developed beamforming techniques combine multiple DMA units to form flexible two-stage spatial filtering systems, where the output of each DMA is fed into a higher-level filter, called virtual filter, for further processing. In this manuscript, we analyze and discuss some properties of a broad class of two-stage beamformers with arbitrary planar geometry. In this context, the DMA units are all assumed to have the same directivity pattern of arbitrary order and can be characterized by a variable number of omnidirectional sensors organized in an arbitrary geometry. For any given choice of the virtual array filter, we introduce a closed-form optimization procedure to design DMA filters that maximize the White Noise Gain (WNG) or the Directivity Factor (DF) of the resulting two-stage beamformer at any frequency. Based on this frequency-dependent design, we propose a frequency-invariant design of the two-stage beamformer and we compare the performance of the two approaches. Finally, we propose two possible computational schemes for the proposed generic two-stage spatial filtering system and discuss their efficiency in performing filtering, steering, and changing beampattern. Davide Albertini, Alberto Bernardini, Federico Borra, Fabio Antonacci, Augusto Sarti |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2022 | On the Prediction of the Frequency Response of a Wooden Plate from Its Mechanical ParametersabstractInspired by deep learning applications in structural mechanics, we focus on how to train two predictors to model the relation between the vibrational response of a prescribed point of a wooden plate and its material properties. In particular, the eigenfrequencies of the plate are estimated via multilinear regression, whereas their amplitude is predicted by a feedforward neural network. We show that labeling the train set by mode numbers instead of by the order of appearance of the eigenfrequencies greatly improves the accuracy of the regression and that the coefficients of the multilinear regressor allow the definition of a linear relation between the first eigenfrequencies of the plate and its material properties. David Giuseppe Badiane, Raffaele Malvermi, Fabio Antonacci, Augusto Sarti |
ICASSP | 5 |
| 2022 | Deepfake Speech Detection Through Emotion Recognition: A Semantic ApproachabstractIn recent years, audio and video deepfake technology has advanced relentlessly, severely impacting people’s reputation and reliability. Several factors have facilitated the growing deepfake threat. On the one hand, the hyper-connected society of social and mass media enables the spread of multimedia content worldwide in real-time, facilitating the dissemination of counterfeit material. On the other hand, neural network-based techniques have made deepfakes easier to produce and difficult to detect, showing that the analysis of low-level features is no longer sufficient for the task. This situation makes it crucial to design systems that allow detecting deepfakes at both video and audio levels. In this paper, we propose a new audio spoofing detection system leveraging emotional features. The rationale behind the proposed method is that audio deepfake techniques cannot correctly synthesize natural emotional behavior. Therefore, we feed our deepfake detector with high-level features obtained from a state-of-the-art Speech Emotion Recognition (SER) system. As the used descriptors capture semantic audio information, the proposed system proves robust in cross-dataset scenarios outperforming the considered baseline on multiple datasets. Emanuele Conti, Davide Salvi, Clara Borrelli, Brian C. Hosler, Paolo Bestagini, Fabio Antonacci, Augusto Sarti, Matthew C. Stamm, Stefano Tubaro |
ICASSP | 7 |
| 2022 | A Data-Driven Approach for Acoustic Parameter Similarity Estimation of Speech RecordingabstractSpeech audio acquisitions exhibit different quality and reverberation properties depending on the recording setup and environment. For this reason, it is expected that speech analysis systems that work correctly on certain audio recordings may fail on others acquired in different acoustic contexts. Therefore, to be able to tell whether a track under analysis shares the same acoustic characteristics of a reference one may be useful to understand if it can be successfully processed by a given speech analysis system. Alternatively, in a forensic scenario, an estimate of acoustic parameter similarity between two tracks can be used to verify whether the recordings have been likely acquired in the same environment or not. In this work, we propose two methods to estimate acoustic parameter similarity between a speech recording under analysis and a reference one. The first method relies on the estimation of channel-based acoustic indicators that are then compared to extract a similarity measure. The second method directly learns a parameter similarity measure through siamese neural networks. Mattia Papa, Clara Borrelli, Paolo Bestagini, Fabio Antonacci, Augusto Sarti, Stefano Tubaro |
ICASSP | 5 |
| 2022 | Sparsity-Based Sound Field Separation in the Spherical Harmonics DomainabstractSound field analysis and reconstruction has been a topic of intense research in the last decades for its multiple applications in spatial audio processing tasks. In this context, the identification of the direct and reverberant sound field components is a problem of great interest, where several solutions exploiting spherical harmonics representations have already been proposed. However, the available techniques demand a large number of high-order microphones (HOMs) and high computational power in order to fulfill the necessary spatial sampling requirements, which can only be reduced by prior information obtained through acoustic measurements. Inspired by compressed sensing approaches, this paper proposes an alternative sparse formulation for estimating the exterior and interior sound field components in the spherical harmonics domain that allows to reduce hardware requirements without the need for additional acoustic measurements. The results show that a considerable reduction in the number of HOMs can be achieved while improving the estimation of the sound field components. Mirco Pezzoli, Maximo Cobos, Fabio Antonacci, Augusto Sarti |
ICASSP | 4 |
| 2022 | A Time-Domain Virtual Bass Enhancement Circuital Model for Real-Time Music ApplicationsabstractIn consumer electronics, the advent of ultra-thin devices has raised interest in Virtual Bass Enhancement (VBE) algorithms for enhancing the acoustic performance of their small-size loudspeakers. In fact, due to physical limitations, large volume velocities cannot be achieved, impairing thus the reproduction of low frequencies. VBE techniques exploit psychoacoustic effects originated by the sound signal processing happening in the inner ear and brain. In this paper, we propose a nonlinear circuital model of a generic time-domain VBE system, and we implement it in the discrete-time domain. As required by most applications of interest, the proposed VBE algorithm is able to operate in real-time. A MUSHRA-like test is then employed to evaluate the bass enhancement performance of the proposed algorithm using different nonlinear devices and parameter configurations. Riccardo Giampiccolo, Alberto Bernardini, Augusto Sarti |
MMSP | 3 |
| 2022 | Group Dictionary Equivalent Source Method for Sparse Nearfield Acoustic HolographyabstractIn this article we propose a novel methodology for the non-invasive estimation of the geometry of the modes of vibration of a vibrating structure, which uses the signals captured by a microphone array placed in close proximity of the vibrating structure. We propose a measurement approach based on an efficient formulation of the Nearfield Acoustic Holography (NAH) problem, using a small sets of equivalent sources that are able to represent a continuous vibrating structure. In this new reformulation of the solution, a Group Dictionary Equivalent Sources Method is developed, which combines the efficiency and flexibility of equivalent sources, with a sparse reconstruction scheme that takes into account prior knowledge about the mechanical behavior of the object under investigation computed through FEA. Riccardo R. De Lucia, Antonio Canclini, Fabio Antonacci, Augusto Sarti |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2021 | Interpolation of Irregularly Sampled Frequency Response Functions Using Convolutional Neural NetworksabstractIn the field of structural mechanics, classical methods for the vibrational characterization of objects exploit the inherent redundancy of a relevant amount of measurements acquired over regular sampling grids. However, there are cases in which parts of the objects under analysis are not accessible with sensors, leading to irregular sampling grids characterized by holes. Recent works have proved the benefits of adding prior knowledge in these scenarios, either through the definition of a suitable decomposition or using Finite Element modelling. In this paper we propose to use Convolutional Autoencoders (CA) for Frequency Response Function (FRF) interpolation from grids with different subsampling schemes. CA learn a compressed representation from a dataset of FRFs synthetized through Finite Element Analysis. Experiments with numerical and experimental data show the effectiveness of the model with a different amount of missing data and its ability to predict real FRFs characterized by different damping and sampling frequency. Matteo Acerbi, Raffaele Malvermi, Mirco Pezzoli, Fabio Antonacci, Augusto Sarti, Roberto Corradi |
ICASSP | 5 |
| 2021 | Arrays of First-Order Steerable Differential MicrophonesabstractThe literature is rich with techniques for the design of small-size Differential Microphone Arrays (DMAs), known for their almost frequency-invariant beampatterns and low computational cost. Few works, instead, discuss the properties of beamformers based on multiple DMA units. In this paper, we consider arbitrarily shaped planar arrays of DMA units. In turn, each DMA unit is a first-order continuously-steerable differential microphone characterized by an arbitrary configuration of omnidirectional sensors and a symmetric beampattern. We present a beamforming technique that, assumed all the DMA units to steer identical beams in the same direction, allows us to approach the behavior of a Delay-And-Sum beamformer or Super-Directive beamformer by solely varying a single scalar parameter. Efficient implementations of the proposed beamformers can be developed by taking into account that, for a wide range of frequencies, the values of such a parameter are practically invariant with respect to the geometry of the array. Federico Borra, Alberto Bernardini, Ivan Bertuletti, Fabio Antonacci, Augusto Sarti |
ICASSP | 5 |
| 2021 | Sparse Recovery Beamforming and Upscaling in the Ray SpaceabstractWe have been exploring the integration of sparse recovery methods into the ray space transform over the past years and now demonstrate the potential and benefits of beamforming and upscaling signals in the integrated ray space and sparse recovery domain. A primary advantage of the ray space approach derives from its robust ability to integrate information from multiple arrays and viewpoints. Nonetheless, for a given viewpoint, the ray space technique requires a dense array that can be divided into sub-arrays enabling the plenacoustic approach to signal processing. In this work, we explore a method to upscale an array beyond the limits imposed by the inter-microphone distances associated with the array and the concomitant spatial aliasing. In other words, sparse recovery enables one to synthesize or interpolate signals corresponding to an array with a greater number of microphones with a smaller inter-microphone distance. A critical issue is whether or not this interpolative synthesis actually improves array signal processing. This work shows that upscaling signals in the integrated ray space and sparse recovery domain can improve both source localization and separation. Shiduo Yu, Craig T. Jin, Fabio Antonacci, Augusto Sarti |
ICASSP | 4 |
| 2021 | Synthetic speech detection through short-term and long-term prediction tracesabstractAbstract Several methods for synthetic audio speech generation have been developed in the literature through the years. With the great technological advances brought by deep learning, many novel synthetic speech techniques achieving incredible realistic results have been recently proposed. As these methods generate convincing fake human voices, they can be used in a malicious way to negatively impact on today’s society (e.g., people impersonation, fake news spreading, opinion formation). For this reason, the ability of detecting whether a speech recording is synthetic or pristine is becoming an urgent necessity. In this work, we develop a synthetic speech detector. This takes as input an audio recording, extracts a series of hand-crafted features motivated by the speech-processing literature, and classify them in either closed-set or open-set. The proposed detector is validated on a publicly available dataset consisting of 17 synthetic speech generation algorithms ranging from old fashioned vocoders to modern deep learning solutions. Results show that the proposed method outperforms recently proposed detectors in the forensics literature. Clara Borrelli, Paolo Bestagini, Fabio Antonacci, Augusto Sarti, Stefano Tubaro |
EURASIP J. Inf. Secur. | 4 |
| 2021 | Reconstructing Speech From CNN EmbeddingsabstractThe complete understanding of the decision-making process of Convolutional Neural Networks (CNNs) is far from being fully reached. Many researchers proposed techniques to interpret what a network actually “learns” from data. Nevertheless many questions still remain unanswered. In this work we study one aspect of this problem by reconstructing speech from the intermediate embeddings computed by a CNNs. Specifically, we consider a pre-trained network that acts as a feature extractor from speech audio. We investigate the possibility of inverting these features, reconstructing the input signals in a black-box scenario, and quantitatively measure the reconstruction quality by measuring the word-error-rate of an off-the-shelf ASR model. Experiments performed using two different CNN architectures trained for six different classification tasks, show that it is possible to reconstruct time-domain speech signals that preserve the semantic content, whenever the embeddings are extracted before the fully connected layers. Luca Comanducci, Paolo Bestagini, Marco Tagliasacchi, Augusto Sarti, Stefano Tubaro |
IEEE Signal Process. Lett. | 4 |
| 2021 | Ray-Space-Based Multichannel Nonnegative Matrix Factorization for Audio Source SeparationabstractNonnegative matrix factorization (NMF) has been traditionally considered a promising approach for audio source separation. While standard NMF is only suited for single-channel mixtures, extensions to consider multi-channel data have been also proposed. Among the most popular alternatives, multichannel NMF (MNMF) and further derivations based on constrained spatial covariance models have been successfully employed to separate multi-microphone convolutive mixtures. This letter proposes a MNMF extension by considering a mixture model with Ray-Space-transformed signals, where magnitude data successfully encodes source locations as frequency-independent linear patterns. We show that the MNMF algorithm can be seamlessly adapted to consider Ray-Space-transformed data, providing competitive results with recent state-of-the-art MNMF algorithms in a number of configurations using real recordings. Mirco Pezzoli, Julio J. Carabias-Orti, Maximo Cobos, Fabio Antonacci, Augusto Sarti |
IEEE Signal Process. Lett. | 5 |
| 2021 | A Wave Digital Newton-Raphson Method for Virtual Analog Modeling of Audio Circuits with Multiple One-Port NonlinearitiesabstractThe digital implementation of a nonlinear audio circuit often employs the Newton-Raphson (NR) method for solving the corresponding system of implicit ordinary differential equations in the discrete-time domain. Although its quadratic convergence speed makes NR attractive for real-time audio applications, quadratic convergence is not always guaranteed, since it depends on initial conditions, and also divergence might occur. For this reason, especially in the context of Virtual Analog modeling, techniques for increasing the robustness of NR are in order. Among the various approaches, the Wave Digital (WD) formalism recently showed potential to rethink traditional circuit simulation methods. In this manuscript, we discuss an original formulation of the NR method in the WD domain for the solution of audio circuits with multiple one-port nonlinearities. We provide an in-depth theoretical analysis of the proposed iterative method and we show how its quadratic convergence strongly depends on the free parameters (called port resistances) introduced when modeling the reference circuit in the WD domain. In particular, we demonstrate that the size of the basin where the WD NR solver can be initialized to converge on a solution with quadratic speed is a function of the free parameters. We also show that by setting each port resistance value as close as possible to the derivative w.r.t. current of the nonlinear element v-i characteristic we keep the basin size large. We finally implement an audio ring modulator circuit with four diodes in order to test the proposed iterative method. Alberto Bernardini, Enrico Bozzo, Federico Fontana, Augusto Sarti |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2021 | Wave Digital Modeling and Implementation of Nonlinear Audio Circuits With NullorsabstractThe nullor is a theoretical two-port element suitable to model several multi-port devices common in audio circuitry, such as ideal operational amplifiers, operational transconductance amplifiers, and transistors operating in linear regime. In this manuscript, we present an approach for the Wave Digital (WD) modeling and implementation of circuits with multiple nullors. In particular, we propose an approach to compute scattering matrices of WD topological junctions absorbing nullors that is less computationally demanding than the techniques available in the literature on WD Filters. We show that the proposed approach turns out to be particularly useful when simulating nonlinear circuits through the Scattering Iterative Method (SIM), a WD fixed-point method recently developed for the solution of circuits with multiple nonlinearities, because it requires a frequent update of the scattering matrices. We also provide a novel convergence analysis of SIM applied to WD structures composed of multiple one-port nonlinear elements and a topological junction absorbing nullors. In order to verify the effectiveness of the proposed methodology, we discuss some WD implementations of analog audio circuits with multiple diodes and opamps, including a precision half-wave rectifier and a wave folder circuit. Riccardo Giampiccolo, Mauro Giuseppe de Bari, Alberto Bernardini, Augusto Sarti |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2021 | Vector Wave Digital Filters and Their Application to Circuits With Two-Port ElementsabstractWave Digital Filters (WDFs) turn circuits into networks of input-output relationships that can be computed in an explicit fashion. This is done through a linear port-wise mapping of Kirchhoff variables into pairs of incident-reflected waves introducing one scalar free parameter per port, called reference port resistance. Parameters are then used to eliminate the implicit equations relating wave variables, referred to as delay-free-loops. Unfortunately, this methodology can only be applied under strong linearity and topological conditions. This manuscript presents an extension of the WDF formalism involving a novel “cross-port” vector definition of waves, whose reference resistance is a matrix of free parameters. This generalization greatly simplifies the WDF implementation of circuits with two-port elements, such as operational amplifiers. It allows us to derive wave-based descriptions of elements such as nullors, for which no scattering relation is available in the literature. Moreover, it enables a full adaptation of a wide class of two-port elements, thus avoiding the delay-free-loops that would otherwise form in traditional WDFs. This new formalism allows us to implement a wider range of circuits with two-port elements in a modular fashion, since the topology and the elements can be modeled independently. Alberto Bernardini, Paolo Maffezzoni, Augusto Sarti |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2020 | Time Difference of Arrival Estimation from Frequency-Sliding Generalized Cross-Correlations Using Convolutional Neural NetworksabstractThe interest in deep learning methods for solving traditional signal processing tasks has been steadily growing in the last years. Time delay estimation (TDE) in adverse scenarios is a challenging problem, where classical approaches based on generalized cross-correlations (GCCs) have been widely used for decades. Recently, the frequency-sliding GCC (FS-GCC) was proposed as a novel technique for TDE based on a sub-band analysis of the cross-power spectrum phase, providing a structured two-dimensional representation of the time delay information contained across different frequency bands. Inspired by deep-learning-based image denoising solutions, we propose in this paper the use of convolutional neural networks (CNNs) to learn the time-delay patterns contained in FS-GCCs extracted in adverse acoustic conditions. Our experiments confirm that the proposed approach provides excellent TDE performance while being able to generalize to different room and sensor setups. Luca Comanducci, Maximo Cobos, Fabio Antonacci, Augusto Sarti |
ICASSP | 4 |
| 2020 | Efficient Implementations of First-Order Steerable Differential Microphone Arrays With Arbitrary Planar GeometryabstractWe present a spatial filtering approach to first-order steerable Differential Microphone Arrays (DMAs) with arbitrary planar geometry. In particular, the design of the spatial filter is based on a recently proposed frequency-domain design methodology that approximates, in a least-square sense, a target beampattern using the Jacobi-Anger expansion involving Bessel functions. Despite the generality of that approach, however, its computational cost turns out to be excessive when working with limited processing resources. The beamforming technique proposed in this manuscript overcomes this issue by exploiting the fact that in DMAs the spacing between sensors is typically smaller than the smallest wavelength of audio signals of interest. This allows us to substitute zero- and first-order Bessel functions with their Taylor series approximation truncated to the first order. Moreover, we show that this approximation allows us to derive an efficient discrete-time-domain implementation of first-order steerable differential beamformers based on arrays with arbitrary geometries. Federico Borra, Alberto Bernardini, Fabio Antonacci, Augusto Sarti |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2020 | A Methodology for the Robust Estimation of the Radiation Pattern of Acoustic SourcesabstractWe propose a novel methodology for estimating the radiation pattern of acoustic sources, which is general enough as to be suitable for a wide variety of sources without the need of anechoic conditions of operation. Multiple plenacoustic cameras (which can be thought of as arrays of acoustic cameras) scan the source while keeping reflections and interferers at bay through deconvolution and windowing of the measured response. In the case of a moving source (e.g. a musical instrument while it is being played), the plenacoustic cameras are also used for tracking the position of the source. As for its orientation, we propose practical solutions for tracking that as well, whenever such information is not known in advance. Two experiments are conducted in order to validate the proposed solution. The former focuses on a commercial loudspeaker cabinet, whose radiation pattern is known in advance and can be used as groundtruth. The latter concerns violins, which exhibit an extremely rich and hard to predict acoustic behavior, due to their inherent structural and constructional complexity. Our method allows us to capture the radiation pattern of the instrument while it is being played, thus returning data corresponding to the natural timbre of the instrument, including the unavoidable acoustic shadow of the violinist's head. Experimental results confirm a relevant improvement in accuracy and robustness afforded by the adoption of dynamic plenacoustic solutions with respect to state-of-the-art techniques. Antonio Canclini, Fabio Antonacci, Stefano Tubaro, Augusto Sarti |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2020 | Frequency-Sliding Generalized Cross-Correlation: A Sub-Band Time Delay Estimation ApproachabstractThe generalized cross-correlation(GCC) is regarded as the most popular approach for estimating the time difference of arrival (TDOA) between the signals received at two sensors. Time delay estimates are obtained by maximizing the GCC output, where the direct-path delay is usually observed as a prominent peak. Moreover, GCCs play also an important role in steered response power (SRP) localization algorithms, where the SRP functional can be written as an accumulation of the GCCs computed from multiple sensor pairs. Unfortunately, the accuracy of TDOA estimates is affected by multiple factors, including noise, reverberation and signal bandwidth. In this paper, a sub-band approach for time delay estimation aimed at improving the performance of the conventional GCC is presented. The proposed method is based on the extraction of multiple GCCs corresponding to different frequency bands of the cross-power spectrum phase in a sliding-window fashion. The major contributions of this paper include: 1) a sub-band GCC representation of the cross-power spectrum phase that, despite having a reduced temporal resolution, provides a more suitable representation for estimating the true TDOA; 2) such matrix representation is shown to be rank one in the ideal noiseless case, a property that is exploited in more adverse scenarios to obtain a more robust and accurate GCC; 3) we propose a set of low-rank approximation alternatives for processing the sub-band GCC matrix, leading to better TDOA estimates and source localization performance. An extensive set of experiments is presented to demonstrate the validity of the proposed approach. Maximo Cobos, Fabio Antonacci, Luca Comanducci, Augusto Sarti |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2020 | Source Localization Using Distributed Microphones in Reverberant Environments Based on Deep Learning and Ray Space TransformabstractIn this article we present a methodology for source localization in reverberant environments from Generalized Cross Correlations (GCCs) computed between spatially distributed individual microphones. Reverberation tends to negatively affect localization based on Time Differences of Arrival (TDOAs), which become inaccurate due to the presence of spurious peaks in the GCC. We therefore adopt a data-driven approach based on a convolutional neural network, which, using the GCCs as input, estimates the source location in two steps. It first computes the Ray Space Transform (RST) from multiple arrays. The RST is a convenient representation of the acoustic rays impinging on the array in a parametric space, called Ray Space. Rays produced by a source are visualized in the RST as patterns, whose position is uniquely related to the source location. The second step consists of estimating the source location through a nonlinear fitting, which estimates the coordinates that best approximate the RST pattern obtained through the first step. It is worth noting that training can be accomplished on simulated data only, thus relaxing the need of actually deploying microphone arrays in the acoustic scene. The localization accuracy of the proposed techniques is similar to the one of SRP-PHAT, however our method demonstrates an increased robustness regarding different distributed microphones configurations. Moreover, the use of the RST as an intermediate representation makes it possible for the network to generalize to data unseen during training. Luca Comanducci, Federico Borra, Paolo Bestagini, Fabio Antonacci, Stefano Tubaro, Augusto Sarti |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2020 | A Parametric Approach to Virtual Miking for Sources of Arbitrary DirectivityabstractIn this article we propose a methodology for the reconstruction of sound fields in arbitrary locations based on the signals acquired by a spatial distribution of compact microphone arrays (virtual miking). The proposed method is suitable for operating in reverberant environments, thanks to a two-stage analysis process, the former of which aims at separating the direct and the diffuse components of the sound field. The method that we propose is inherently parametric, as the sources of the acoustic scene are characterized by parameters describing location and directivity (spherical harmonics expansion), which are extracted from the exterior model of the direct component of the sound field. Once the parameters of the sources are extracted, the direct sound field at an arbitrary location is reconstructed. The diffuse component is reconstructed from the joint knowledge of the diffuse component at the locations of the distributed microphone arrays, under the assumption of isotropic behavior. Results show that the proposed technique is able to analyze the sound field and reconstruct the parameters of the sources that are active in the scene. In addition, the synthesis of the signals at the virtual microphone locations turns out to accurately match (in terms of spatial cues) the actual sound field, as measured by a microphone places in the desired location. Mirco Pezzoli, Federico Borra, Fabio Antonacci, Stefano Tubaro, Augusto Sarti |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2020 | 3D Room Geometry Inference Using a Linear Loudspeaker Array and a Single MicrophoneabstractSound reproduction systems may highly benefit from detailed knowledge of the acoustic space to enhance the spatial sound experience. This article presents a room geometry inference method based on identification of reflective boundaries using a high-resolution direction-of-arrival map produced via room impulse responses (RIRs) measured with a linear loudspeaker array and a single microphone. Exploiting the sparse nature of the early part of the RIRs, Elastic Net regularization is applied to obtain a 2D polar-coordinate map, on which the direct path and early reflections appear as distinct peaks, described by their propagation distance and direction of arrival. Assuming a separable room geometry with four side-walls perpendicular to the floor and ceiling, and imposing pre-defined geometrical constraints on the walls, the 2D-map is segmented into six regions, each corresponding to a particular wall. The salient peaks within each region are selected as candidates for the first-order wall reflections, and a set of potential room geometries is formed by considering all possible combinations of the associated peaks. The room geometry is then inferred using a cost function evaluated on the higher-order reflections computed via beam tracing. The proposed method is tested with both simulated and measured data. Cagdas Tuna, Antonio Canclini, Federico Borra, Philipp Götz, Fabio Antonacci, Andreas Walther 0001, Augusto Sarti, Emanuël A. P. Habets |
IEEE ACM Trans. Audio Speech Lang. Process. | 7 |
| 2019 | Linear Multistep Discretization Methods With Variable Step-Size in Nonlinear Wave Digital Structures for Virtual Analog ModelingabstractThere is a growing interest in Virtual Analog modeling algorithms for musical audio processing designed in the Wave Digital (WD) domain. Such algorithms typically employ a discretization strategy based on the trapezoidal rule with fixed sampling step, though this is not the only option. In fact, alternative discretization strategies (possibly with an adaptive sampling step) can be quite advantageous, particularly when dealing with nonlinear systems characterized by stiff equations. In this paper, we propose a unified approach for modeling capacitors and inductors in the WD domain using generic linear multi-step discretization methods with variable time-step size, and provide generalized adaptation conditions. We also show that the proposed approach for implementing dynamic (energy-storing) elements in the WD domain is particularly suitable to be combined with a recently developed technique for efficiently solving a class of circuits with multiple one-port nonlinearities, called Scattering Iterative Method. Finally, as examples of application, we develop WD models for a Van Der Pol oscillator and a dynamic diode-based ring modulator, which use different discretization methods. Alberto Bernardini, Paolo Maffezzoni, Augusto Sarti |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2019 | Uniform Linear Arrays of First-Order Steerable Differential MicrophonesabstractWe propose a spatial filtering method for linear arrays of First-Order Steerable Differential Microphones (FOSDMs), which operates in two layers. In the former, signals acquired by individual microphones are locally filtered to produce the outputs of the FOSDMs. In the latter, the outputs of the FOSDMs are processed by another filter. We analyse different design methodologies and study the conditions under which the two filtering layers can be decoupled. The proposed two-layer spatial filter can be flexibly controlled with a single scalar parameter, which can be chosen, for example, to maximize the White Noise Gain (like in a Delay-and-Sum beamformer); or to maximize the Directivity Factor (like in a Super-Directive beamformer); without needing any matrix inversion. The effectiveness of the proposed beamforming method is compared with traditional spatial filtering techniques using different metrics. Federico Borra, Alberto Bernardini, Fabio Antonacci, Augusto Sarti |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2019 | A Presence- and Performance-Driven Framework to Investigate Interactive Networked Music Learning ScenariosabstractCooperative music making in networked environments has been subject of extensive research, scientific and artistic. Networked music performance (NMP) is attracting renewed interest thanks to the growing availability of effective technology and tools for computer-based communications, especially in the area of distance and blended learning applications. We propose a conceptual framework for NMP research and design in the context of classical chamber music practice and learning: presence-related constructs and objective quality metrics are used to problematize and systematize the many factors affecting the experience of studying and practicing music in a networked environment. To this end, a preliminary NMP experiment on the effect of latency on chamber music duos experience and quality of the performance is introduced. The degree of involvement, perceived coherence, and immersion of the NMP environment are here combined with measures on the networked performance, including tempo trends and misalignments from the shared score. Early results on the impact of temporal factors on NMP musical interaction are outlined, and their methodological implications for the design of pedagogical applications are discussed. Stefano Delle Monache, Luca Comanducci, Michele Buccoli, Massimiliano Zanoni, Augusto Sarti, Enrico Pietrocola, Filippo Berbenni, Giovanni Cospito |
Wirel. Commun. Mob. Comput. | 5 |
| 2018 | A Weighted Least Squares Beam Shaping Technique for Sound Field ControlabstractA weighted least squares beam shaping technique for sound field control using a loudspeaker array is proposed. Given a desired spatial response at prescribed control points, the space-time filter is designed by solving a least squares minimization problem. To reduce the computational effort, we propose to place control points only along an arc of circumference centered at the center of the array and passing through a region of interest. Furthermore, we adopt a weighted least squares approach for the design of the space-time filter, so that control points at directions towards which we admit a looser control of the sound field are less relevant in the filter design. The choice of the weights depends on the specific application and we demonstrate the feasibility of the proposed approach for sound zones scenario with one bright and one dark zone. Antonio Canclini, Dejan Markovic, Martin Schneider 0009, Fabio Antonacci, Emanuël A. P. Habets, Andreas Walther 0001, Augusto Sarti |
ICASSP | 7 |
| 2018 | Considerations Regarding Individualization of Head-Related Transfer FunctionsabstractThis paper provides some considerations regarding using individualized head-related transfer functions for rendering binaural spatial audio over headphones. It briefly considers the degree of benefit that individualization may provide. It then examines the degree of variation existing within the ear morphology across listeners within the Sydney-York Morphological and Recording of Ears (SYMARE) database using kernel principal component analysis and the large deformation diffeomorphic metric mapping framework. The degree of variation across listeners in the directivity patterns associated with head-related transfer functions is also analyzed as a function of frequency. The variation in ear morphology is related to the variation in the directivity patterns using simple linear regression. Craig T. Jin, Reza Zolfaghari, Xian Long, Arun Sebastian, Shayikh Hossain, Joan Glaunès, Anthony I. Tew, Muhammad Shahnawaz, Augusto Sarti |
ICASSP | 9 |
| 2018 | Estimation of the Sound Field at Arbitrary Positions in Distributed Microphone Networks Based on Distributed Ray Space TransformabstractIn this paper we propose a parametric sound field reconstruction approach. In particular, the technique is based on the estimation of three parameters for each acoustic source (source position, radiation pattern and source signal) given the signals acquired by few arbitrarily placed microphone arrays. This allows us to synthesize the signal of a virtual microphone placed in any point of the acoustic scene. Mirco Pezzoli, Federico Borra, Fabio Antonacci, Augusto Sarti, Stefano Tubaro |
ICASSP | 4 |
| 2018 | Wave Digital-Based Variability Analysis of Electrical Mismatch in Photovoltaic ArraysabstractThis research investigates the effects that electrical mismatches and partial shading can have on the performance of photovoltaic arrays. The analysis adopts a probabilistic point of view where the most relevant parameters of the solar units are seen as random variables. The analysis relies on an efficient and robust simulation technique, based on Wave Digital principles, that is tailored to the modular topology of solar arrays. It is shown how electrical mismatch, solar shading and array topology can interact among them in a quite complex way. Alberto Bernardini, Augusto Sarti, Paolo Maffezzoni, Luca Daniel |
ISCAS | 2 |
| 2018 | Wave Digital Implementation of Robust First-Order Differential Microphone ArraysabstractIn this letter, a novel time-domain implementation of robust first-order differential microphone arrays (DMAs), based on wave digital filters, is presented. The proposed beamforming method is extremely efficient, as it requires at most two multipliers and one delay for each filter, where the necessary number of filters equals the number of physical microphones of the array, and it avoids the use of fractional delays. The update of the coefficients of the filters, required for reshaping the beampattern, has a significantly lower computational cost with respect to the time-domain methods presented in the literature. This makes the proposed method suitable for real-time DMA applications with time-varying beampatterns. Alberto Bernardini, Fabio Antonacci, Augusto Sarti |
IEEE Signal Process. Lett. | 3 |
| 2017 | Dictionary-based Equivalent Source Method for Near-Field Acoustic HolographyabstractIn this paper, we propose a modification of the standard Equivalent Source Method (ESM) for Near-Field Acoustic Holography (NAH). As in EMS, we aim at modeling the acoustic pressure radiated from a vibrating object, and its surface velocity, as the joint effect of a set of equivalent sources located within or close to the object itself. The estimation of the equivalent source strengths (weigths) comes from the solution of a highly ill-conditioned problem. Rather than solving this problem in the least-squares sense, we exploit the 3D model of the vibrating object, along with a rough estimate of its physical parameters, to restrict the space of the solutions. More specifically, we make use of Finite Element Analysis for populating a compressed dictionary of possible equivalent source weights. NAH is then approached by seeking a sparse linear combination of the entries of the dictionary. Experiments carried on a public database prove the effectiveness of the proposed technique, especially when the number of available microphones is limited, and in the presence of a significant level of measurement noise. Antonio Canclini, Massimo Varini, Fabio Antonacci, Augusto Sarti |
ICASSP | 4 |
| 2017 | Modeling Sallen-Key audio filters in the Wave Digital domainabstractSallen-Key filters are widespread in audio circuits. Therefore, accurate and efficient digital models of such filters are highly desirable in audio Virtual Analog applications. In this paper, we will discuss a possible strategy, based on Wave Digital Filters (WDFs), for implementing all the analog filters described in the historical 1955 manuscript by Sallen and Key. In particular, we will group the eighteen filter models presented by Sallen and Key into nine classes, according to their topological properties. For each class we will describe the corresponding WDF structures. Finally, we will compare the output signals of WDFs to the output signals of the same models implemented in LTSpice. Mattia Verasani, Alberto Bernardini, Augusto Sarti |
ICASSP | 3 |
| 2017 | Using multi-dimensional correlation for matching and alignment of MoCap and video signalsabstractMotion analysis and tracking often relies on multimodal signals, e.g., video, depth map, motion capture (MoCap), due to the completeness of information they jointly provide. The joint analysis of multimodal signals requires to know the correct timing, i.e., the signals to be aligned. In this paper we propose an approach to automatically estimate the correct matching and alignment between a video and a MoCap recording acquired from the same session, based on the multi-dimensional correlation of velocity-based features extracted from the two recordings. We validate our approach over a dataset of dance recordings of four genres, and we achieve promising results for both the alignment and matching scenarios. Michele Buccoli, Bruno Di Giorgi, Massimiliano Zanoni, Fabio Antonacci, Augusto Sarti |
MMSP | 5 |
| 2017 | Multicamera rig calibration by double-sided thick checkerboardabstractA multi‐camera rig calibration algorithm based on a double sided planar target is proposed. Due to their inherently simple realisation, low cost and accuracy, planar calibration targets came out as one of the most largely adopted calibration tools both for intrinsic and extrinsic camera parameters. However, concerning the estimation of extrinsic parameters, one of the major drawbacks of these targets is their requirement for distinct target visibility from both cameras. This prevents many configurations from being adopted where, e.g. two cameras are facing each other. An inexpensive solution could be based on printing/pasting a planar pattern on both target sides, however, the relative misalignment between the patterns on the two sides and the target thickness could be unknown. The authors propose a solution where double‐sided target displacement error is estimated together with the extrinsic parameters allowing the reuse of all the available planar calibration tools in less constrained configurations. To assess their approach the authors tested the system in two scenarios, one using two professional 4K cameras and one using two smartphones. Marco Marcon, Augusto Sarti, Stefano Tubaro |
IET Comput. Vis. | 2 |
| 2017 | Efficient Continuous Beam Steering for Planar Arrays of Differential MicrophonesabstractPerforming continuous beam steering, from planar arrays of high-order differential microphones, is not trivial. The main problem is that shape-preserving beams can be steered only in a finite set of privileged directions, which depend on the position and the number of physical microphones. In this letter, we propose a simple and computationally inexpensive method for alleviating this problem using planar microphone arrays. Given two identical reference beams pointing in two different directions, we show how to build a beam of nearly constant shape, which can be continuously steered between such two directions. The proposed method, unlike the diffused steering approaches based on linear combinations of eigenbeams (spherical harmonics), is applicable to planar arrays also if we deal with beams characterized by high-order polar patterns. Using the coefficients of the Fourier series of the polar patterns, we also show how to find a tradeoff between shape invariance of the steered beam, and maximum angular displacement between the two reference beams. We show the effectiveness of the proposed method through the analysis of models based on first-, second-, and third-order differential microphones. Alberto Bernardini, Matteo D'Aria, Roberto Sannino, Augusto Sarti |
IEEE Signal Process. Lett. | 4 |
| 2017 | A Data-Driven Model of Tonal Chord Sequence ComplexityabstractWe present a compound language model of tonal chord sequences, and evaluate its capability to estimate perceived harmonic complexity. In order to build the compound model, we trained three different models: prediction by partial matching, a hidden Markov model and a deep recurrent neural network on a novel large dataset containing half a million annotated chord sequences. We describe the training process and propose an interpretation of the harmonic patterns that are learned by the hidden states of these models. We use the compound model to generate new chord sequences and estimate their probability, which we then relate to perceived harmonic complexity. In order to collect subjective ratings of complexity, we devised a listening test comprising two different experiments. In the first, subjects choose the more complex chord sequence between two. In the second, subjects rate with a continuous scale the complexity of a single chord sequence. The results of both experiments show a strong relation between negative log probability, given by our language model, and the perceived complexity ratings. The relation is stronger for subjects with high musical sophistication index, acquired through the GoldMSI standard questionnaire. The analysis of the results also includes the preference ratings that have been collected along with the complexity ratings; a weak negative correlation emerged between preference and log probability. Bruno Di Giorgi, Simon Dixon, Massimiliano Zanoni, Augusto Sarti |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2017 | Distributed 3D Source Localization from 2D DOA Measurements Using Multiple Linear ArraysabstractThis manuscript addresses the problem of 3D source localization from direction of arrivals (DOAs) in wireless acoustic sensor networks. In this context, multiple sensors measure the DOA of the source, and a central node combines the measurements to yield the source location estimate. Traditional approaches require 3D DOA measurements; that is, each sensor estimates the azimuth and elevation of the source by means of a microphone array, typically in a planar or spherical configuration. The proposed methodology aims at reducing the hardware and computational costs by combining measurements related to 2D DOAs estimated from linear arrays arbitrarily displaced in the 3D space. Each sensor measures the DOA in the plane containing the array and the source. Measurements are then translated into an equivalent planar geometry, in which a set of coplanar equivalent arrays observe the source preserving the original DOAs. This formulation is exploited to define a cost function, whose minimization leads to the source location estimation. An extensive simulation campaign validates the proposed approach and compares its accuracy with state-of-the-art methodologies. Antonio Canclini, Fabio Antonacci, Augusto Sarti, Stefano Tubaro |
Wirel. Commun. Mob. Comput. | 3 |
| 2016 | A linear operator for the computation of soundfield mapsabstractIn the process of soundfield imaging, as defined in the literature, a microphone array is subdivided into overlapping sub-arrays and soundfield images are obtained by juxtaposition of spatial spectra computed from individual subarray data. In this paper we show that the whole process can be conveniently seen as a linear transformation applied to array data. This linear transformation embeds a nonlinear mapping to cast the directional information in a more convenient domain: the ray space. We show by simulations that the proposed formulation is suitable for fast implementation of the soundfield imaging operation, and, more specifically, for the localization of acoustic sources. Lucio Bianchi, V. Baldini Anastasio, Dejan Markovic, Fabio Antonacci, Augusto Sarti, Stefano Tubaro |
ICASSP | 5 |
| 2016 | A low-cost solution to 3D pinna modeling for HRTF predictionabstractWe propose an infrared (IR) stereo-vision system for estimating the 3D model of the pinna, based on low-cost devices. A commercial IR calibrated stereo camera is used in conjunction with a structured IR light projector, to acquire highly textured snapshots of the pinna. A point cloud is computed for each snapshot by triangulating the stereo correspondences detected in the acquired IR images. A complete 3D model is computed by aligning and merging the point clouds, and then creating a polygonal mesh surface. The nominal accuracy of the proposed system turns to be about 1 mm, which enables an accurate prediction of the Head Related Transfer Function (HRTF) through numerical acoustic simulation. Luca Bonacina, Antonio Canclini, Fabio Antonacci, Marco Marcon, Augusto Sarti, Stefano Tubaro |
ICASSP | 5 |
| 2016 | Toothbrush motion analysis to help children learn proper tooth brushing
Marco Marcon, Augusto Sarti, Stefano Tubaro |
Comput. Vis. Image Underst. | 2 |
| 2016 | Extraction of Acoustic Sources Through the Processing of Sound Field Maps in the Ray SpaceabstractOur goal is to develop a model-based approach to acoustic source extraction from microphone array data, which is suitable for both near-field and far-field sources. A signal representation based on plane-wave (PW) decomposition is suitable for acoustic sources in the far field as the resulting spectrum turns out to be impulsive. When the source approaches the array, however, the curvature of the wavefront causes the spectrum of the PW components to depart from impulsive behavior, thus making source extraction harder to attain. In this paper, we adopt a sound field representation based on the local estimation of the plenacoustic function along the array line. This approach consists of dividing the array into subarrays, and applying the PW analysis on individual subarrays. This has the immediate result of extending the range of validity of the far-field hypothesis, as a source that enters the near-field range of the extended array is still in the far-field range of the subarrays. PW analysis on subarrays allows us to construct the so-called sound field map in a domain of acoustic visibility called ray space. The extraction of the desired source is accomplished through spatial filtering of the sound field map. The design of the spatial filter relies on a linear minimum mean square error criterion defined on the sound field map. The effectiveness of the proposed methodology is proven through an extensive simulation campaign as well as real experiments. Dejan Markovic, Fabio Antonacci, Lucio Bianchi, Stefano Tubaro, Augusto Sarti |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2016 | 3D Beam Tracing Based on Visibility Lookup for Interactive Acoustic ModelingabstractWe present a method for accelerating the computation of specular reflections in complex 3D enclosures, based on acoustic beam tracing. Our method constructs the beam tree on the fly through an iterative lookup process of a precomputed data structure that collects the information on the exact mutual visibility among all reflectors in the environment (region-to-region visibility). This information is encoded in the form of visibility regions that are conveniently represented in the space of acoustic rays using the Plücker coordinates. During the beam tracing phase, the visibility of the environment from the source position (the beam tree) is evaluated by traversing the precomputed visibility data structure and testing the presence of beams inside the visibility regions. The Plücker parameterization simplifies this procedure and reduces its computational burden, as it turns out to be an iterative intersection of linear subspaces. Similarly, during the path determination phase, acoustic paths are found by testing their presence within the nodes of the beam tree data structure. The simulations show that, with an average computation time per beam in the order of a dozen of microseconds, the proposed method can compute a large number of beams at rates suitable for interactive applications with moving sources and receivers. Dejan Markovic, Fabio Antonacci, Augusto Sarti, Stefano Tubaro |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2015 | A Dimensional Contextual Semantic Model for music description and retrievalabstractSeveral paradigms for high-level music descriptions have been proposed to develop effective system for browsing and retrieving musical content in large repositories. Such paradigms are based on either categorical or dimensional models. The interest in dimensional models has recently grown a great deal, as they define a semantic relation between concepts through graded descriptions. One problem that affects semantic descriptions is the ambiguity that often arises from using the same descriptor in different contexts. In order to overcome this difficulty, it is important to model and address polysemy, which is the property of words to take on different meanings depending on the use-context. In this paper we propose a Dimensional Contextual Semantic Model for defining semantic relations among descriptors in a context-aware fashion. This model is here used for developing a semantic music search engine. In order to evaluate the effectiveness of our model, we compare this engine with two systems that are based on different description models. Michele Buccoli, Massimiliano Zanoni, Augusto Sarti, Stefano Tubaro |
ICASSP | 4 |
| 2015 | Piecewise distortion correction for fisheye lensesabstractLens distortion is a well-known problem for camera calibration. In particular in applications where a large amount of low-quality acquisition devices is adopted, like, e.g. embedded systems or Cyber-Physical Systems (CPS), a complete re-sectioning and undistortion in different operating conditions (e.g. autofocus, zooming) could not be feasible and accurate. Usually undistortion is obtained building a proper invertible geometrical distortion model with a specific number of parameters, but, unfortunately with the actually available low cost wide-angle and ultra wide-angle lenses a simple mathematical model for a global closed form solution can be ineffective in practical cases in particular in the peripheral image regions. In order to account for these problems we present a novel local correction approach based on the planarity and orthogonality constrains for a planar target (a checkerboard) where a Look-Up Table (LUT) is built to provide the minimum displacement for every pixel in the target region minimizing interpolation and residual distortion. The proposed method can also be considered a preliminary step in order to recognize image elements before global undistortion e.g. for features matching in images stitching. Marco Marcon, Augusto Sarti, Stefano Tubaro |
ICIP | 2 |
| 2015 | A Robust and Low-Complexity Source Localization Algorithm for Asynchronous Distributed Microphone NetworksabstractIn this paper, we propose a robust and low-complexity acoustic source localization technique based on time differences of arrival (TDOA), which addresses the scenario of distributed sensor networks in 3D environments. Network nodes are assumed to be unsynchronized, i.e., TDOAs between microphones belonging to different nodes are not available. We begin with showing how to select feasible TDOAs for each sensor node, exploiting both geometrical considerations and a characterization of the overall generalized cross correlation (GCC) shape. We then show how to localize sources in the space-range reference frame, where TDOA measurements have a clear geometrical interpretation that can be fruitfully used in the scenario of unsynchronized sensors. In this framework, in fact, the source corresponds to the apex of a hypercone passing through points described by the sole microphone positions and TDOA measurements. The localization problem is therefore approached as a hypercone fitting problem. Finally, in order to improve the robustness of the estimate, we include an outlier detection procedure based on the evaluation of the hypercone fitting residuals. A refinement of source location estimate is then performed ignoring the contributions coming from outlier measurements. A set of simulations shows the performance of individual blocks of the system, with particular focus on the effect of TDOA selection on source localization and refinement steps. Experiments on real data validate the localization algorithm in an everyday scenario, proving that good accuracy can be obtained while saving computational cost in comparison with state-of-the-art techniques. Antonio Canclini, Paolo Bestagini, Fabio Antonacci, Marco Compagnoni, Augusto Sarti, Stefano Tubaro |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2015 | Multiview Soundfield Imaging in the Projective Ray SpaceabstractA soundfield image is a data structure that efficiently encodes and represents the wave field as captured by a microphone array. Its representation is based on the directional plenacoustic function, which is defined as the radiance of the acoustic paths (rays) that cross the segment that the array lies upon. The soundfield image can be processed “as is” to develop a variety of applications. In its original formulation, the soundfield image is based on a Euclidean parameterization that can accommodate a limited range of rays and is suitable for managing a single array only. In this paper, we generalize this methodology to the case of multiple microphone arrays deployed in space. The use of multiple arrays allows us to capture truly global information on the sound field but requires us to rethink the ray space, and adopt a global representation of the acoustic rays based on projective geometry. After introducing the new parameterization, we present two examples of applications: the estimation of the mutual poses of two or more arrays (self-calibration); and the localization of multiple acoustic sources. The effectiveness of these applications is proven through simulations as well as real data experiments. Dejan Markovic, Fabio Antonacci, Augusto Sarti, Stefano Tubaro |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2014 | Robust beamforming under uncertainties in the loudspeakers directivity patternabstractIn this paper we propose a robust beamforming technique which takes into account uncertainties and variations in the radiation pattern of the loudspeakers. The proposed technique is based on the solution of a robust least-square problem in which the propagation matrix is to some extent unknown. Both simulations and experimental results prove the validity of the proposed methodology in terms of directivity index and white noise gain. Lucio Bianchi, R. Magalotti, Fabio Antonacci, Augusto Sarti, Stefano Tubaro |
ICASSP | 4 |
| 2014 | Estimation of Acoustic Reflection Coefficients Through Pseudospectrum MatchingabstractEstimating the geometric and reflective properties of the environment is important for a wide range of applications of space-time audio processing, from acoustic scene analysis to room equalization and spatial audio rendering. In this manuscript, we propose a methodology for frequency-subband in-situ estimation of the reflection coefficients of planar surfaces. This is a rather challenging task, as the reflection coefficients depend on the frequency and the angle of incidence and their estimate is highly sensitive to background noise and interfering sources. Our method is based on the assumption that we know the geometry of the reflectors; the position and the radiation pattern of the source; the position and the spatial response of the array. Applying beamforming algorithms on a single set of measured sensor data, we estimate the angular distribution of the acoustic energy (angular pseudospectrum) that impinges on a microphone array. We then apply a two-step iterative estimation technique based on an Expectation-Maximization (EM) algorithm. The first step estimates the scaling factors. The second one infers the reflection coefficients from the scaling factors. Under the assumption of additive white Gaussian noise, we finally determine the reflection coefficients with a Maximum Likelihood (ML) estimation method. The effectiveness and the accuracy of the proposed technique are assessed through experiments based on measured data. Dejan Markovic, Konrad Kowalczyk, Fabio Antonacci, Christian Hofmann 0001, Augusto Sarti, Walter Kellermann |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2013 | Localization of virtual acoustic sources based on the Hough transform for sound field rendering applicationsabstractIn this paper we propose a methodology for the localization of virtual acoustic sources for sound field rendering applications. After the reconstruction of the sound field in the listening area by means of circular harmonic decomposition, the virtual source location is found through the Hough transform. We prove the accuracy of the proposed methodology by comparing the source locations estimates with those of a subjective test campaign. Lucio Bianchi, Fabio Antonacci, Antonio Canclini, Augusto Sarti, Stefano Tubaro |
ICASSP | 4 |
| 2013 | Rendering of directional sources through loudspeaker arrays based on plane wave decompositionabstractIn this paper we present a technique for the rendering of directional sources by means of loudspeaker arrays. The proposed methodology is based on a decomposition of the sound field in terms of plane waves. Within this framework the directivity of the source is naturally included in the rendering problem, therefore accommodating the directivity into the picture becomes much simpler. For this purpose, the loudspeaker array is subdivided into overlapping sub-arrays, each generating a plane wave component. The individual plane waves are then weighed by the desired directivity pattern. Simulations and experimental results show that the proposed technique is able to reproduce the sound field of directional sources with an improved accuracy with respect to existing techniques. Lucio Bianchi, Fabio Antonacci, Augusto Sarti, Stefano Tubaro |
MMSP | 3 |
| 2013 | A music search engine based on semantic text-based queryabstractSearch and retrieval of songs from a large music repository usually relies on added meta-information (e.g., title, artist or musical genre); or on specific descriptors (e.g. mood); or on categorical music descriptors; none of which can specify the desired intensity. In this work, we propose an early example of semantic text-based music search engine. The semantic description takes into account emotional and non-emotional musical aspects. The method also includes a query-by-similarity search approach performed using semantic cues. We model both concepts and musical content in dimensional spaces that are suitable for carrying intensity information on the descriptors. We process the semantic query with a Natural Language parser to capture only the relevant words and qualifiers. We rely on Bayesian Decision theory to model concepts and songs as probability distributions. The resulted ranked list of songs are produced through a posterior probability model. A prototype of the system has been proposed to 53 subjects for evaluation, with good ratings on performance, usefulness and potential. Michele Buccoli, Massimiliano Zanoni, Augusto Sarti, Stefano Tubaro |
MMSP | 3 |
| 2013 | Acoustic Source Localization With Distributed Asynchronous Microphone NetworksabstractWe propose a method for localizing an acoustic source with distributed microphone networks. Time Differences of Arrival (TDOAs) of signals pertaining the same sensor are estimated through Generalized Cross-Correlation. After a TDOA filtering stage that discards measurements that are potentially unreliable, source localization is performed by minimizing a fourth-order polynomial that combines hyperbolic constraints from multiple sensors. The algorithm turns to exhibit a significantly lower computational cost compared with state-of-the-art techniques, while retaining an excellent localization accuracy in fairly reverberant conditions. Antonio Canclini, Fabio Antonacci, Augusto Sarti, Stefano Tubaro |
IEEE Trans. Speech Audio Process. | 3 |
| 2013 | Soundfield Imaging in the Ray SpaceabstractIn this work we propose a general approach to acoustic scene analysis based on a novel data structure (ray-space image) that encodes the directional plenacoustic function over a line segment (Observation Window, OW). We define and describe a system for acquiring a ray-space image using a microphone array and refer to it as ray-space (or “soundfield”) camera. The method consists of acquiring the pseudo-spectra corresponding to a grid of sampling points over the OW, and remapping them onto the ray space, which parameterizes acoustic paths crossing the OW. The resulting ray-space image displays the information gathered by the sensors in such a way that the elements of the acoustic scene (sources and reflectors) will be easy to discern, recognize and extract. The key advantage of this method is that ray-space images, irrespective of the application, are generated by a common (and highly parallelizable) processing layer, and can be processed using methods coming from the extensive literature of pattern analysis. After defining the ideal ray-space image in terms of the directional plenacoustic function, we show how to acquire it using a microphone array. We also discuss resolution and aliasing issues and show two simple examples of applications of ray-space imaging. Dejan Markovic, Fabio Antonacci, Augusto Sarti, Stefano Tubaro |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2012 | Correction method for nonideal iris recognitionabstractThe use of iris as biometric trait has emerged as one of the most preferred method because of its uniqueness, lifetime stability and regular shape. Moreover it shows public acceptance and new user-friendly capture devices are developed and used in a broadened range of applications. Currently, iris recognition systems work well with frontal iris images from cooperative users. Nonideal iris images are still a challenge for iris recognition and can significantly affect the accuracy of iris recognition systems. In this paper, we propose a method to correct off-angle iris image. Taking into account the eye morphology and the reflectance properties of the external transparent layers, we can evaluate the distorting effect that is present in the acquired image. The correction algorithm proposed includes a first modeling phase of the human eye, a segmentation of the acquired image, and a simulation phase where the acquisition geometry is reproduced and the distortions are evaluated. Finally we obtain an image which does not contain the distorting effects due to jumps in the refractive index. We show how this correction process reduce the intra-class variations for off-angle iris images. Eliana Frigerio, Marco Marcon, Augusto Sarti, Stefano Tubaro |
ICIP | 3 |
| 2012 | 3D wide baseline correspondences using depth-maps
Marco Marcon, Eliana Frigerio, Augusto Sarti, Stefano Tubaro |
Signal Process. Image Commun. | 3 |
| 2012 | Inference of Room Geometry From Acoustic Impulse ResponsesabstractAcoustic scene reconstruction is a process that aims to infer characteristics of the environment from acoustic measurements. We investigate the problem of locating planar reflectors in rooms, such as walls and furniture, from signals obtained using distributed microphones. Specifically, localization of multiple two- dimensional (2-D) reflectors is achieved by estimation of the time of arrival (TOA) of reflected signals by analysis of acoustic impulse responses (AIRs). The estimated TOAs are converted into elliptical constraints about the location of the line reflector, which is then localized by combining multiple constraints. When multiple walls are present in the acoustic scene, an ambiguity problem arises, which we show can be addressed using the Hough transform. Additionally, the Hough transform significantly improves the robustness of the estimation for noisy measurements. The proposed approach is evaluated using simulated rooms under a variety of different controlled conditions where the floor and ceiling are perfectly absorbing. Results using AIRs measured in a real environment are also given. Additionally, results showing the robustness to additive noise in the TOA information are presented, with particular reference to the improvement achieved through the use of the Hough transform. Fabio Antonacci, Jason Filos, Mark R. P. Thomas, Emanuël A. P. Habets, Augusto Sarti, Patrick A. Naylor, Stefano Tubaro |
IEEE Trans. Speech Audio Process. | 5 |
| 2012 | Localization of Acoustic Sources Through the Fitting of Propagation Cones Using Multiple Independent ArraysabstractIn this paper, we propose a novel acoustic source localization method that accommodates the general scenario of multiple independent microphone arrays. The method is based on a 3-D parameter space defined by the 2-D spatial location of a source and the range difference extracted from the time difference of arrival (TDOA). In this space, the set of points that correspond to a given range lie on a circle that expand as the range increases, forming a cone whose apex is the actual location of the source. In this parameter space, the lack of synchronization between arrays results in the fact that clusters of data associated to individual arrays are free to shift along the range axis. The cone constraint, in fact, enables the realignment of such clusters while positioning the cone vertex (source location), thus resulting in a joint data re-synchronization and source localization. We also propose a novel and general analysis methodology for swiftly assessing the localization error as a function of the TDOA uncertainties, which is remarkably accurate for small localization bias. With the aid of this method, simulations and experiments on real data, we show that the cone-fitting process offers excellent localization accuracy in the scenario of multiple unsynchronized arrays, as well as in simpler single-array scenarios, also in comparison with state-of-the-art techniques. We also show that the proposed method offers the desired flexibility for adapting to arbitrary geometries of microphone clusters. Marco Compagnoni, Paolo Bestagini, Fabio Antonacci, Augusto Sarti, Stefano Tubaro |
IEEE Trans. Speech Audio Process. | 4 |
| 2011 | A methodology for evaluating the accuracy of wave field rendering techniquesabstractIn this paper we propose a methodology for assessing the accuracy of techniques of wave field rendering through loudspeaker arrays. In order to measure the rendered wave field we adopt a solution based on a circular harmonic analysis of the sound field captured by a virtual microphone array. As a result of this analysis stage, we are able to compare the target, the theoretical and the measured wave fields, which may differ due to the non-ideality in the loudspeaker array or in the environment that generates some spurious reverberations. Moreover, in order to quantify the error between target, theoretical and measured wave fields, we define some evaluation metrics, based on RMSE and modal analysis of the acquired wave fields. We show some experimental results on real data. Antonio Canclini, Paolo Annibale, Fabio Antonacci, Augusto Sarti, Rudolf Rabenstein, Stefano Tubaro |
ICASSP | 4 |
| 2011 | From direction of arrival estimates to localization of planar reflectors in a two dimensional geometryabstractIn this paper we propose a novel technique to localize planar obstacles through the measurement of the Direction of Arrival by a microphone array. The measurement of the Direction of Arrival of the reflected path is turned into a quadratic constraint where the unknowns are the line parameters of the reflector. A cost function that combines multiple constraints is then derived. A parametric description of the obstacle is found by minimization of the cost function. Some simulations and experimental results show the feasibility of the proposed method. Antonio Canclini, Paolo Annibale, Fabio Antonacci, Augusto Sarti, Rudolf Rabenstein, Stefano Tubaro |
ICASSP | 4 |
| 2011 | A system for dynamic playlist generation driven by multimodal control signals and descriptorsabstractThis work describes a general approach to multimedia playlist generation and description and an application of the approach to music information retrieval. The example of system that we implemented updates a musical playlist on the fly based on prior information (musical preferences); current descriptors of the song that is being played; and fine-grained and semantically rich descriptors (descriptors of user's gestures, of environment conditions, etc.). The system incorporates a learning system that infers the user's preferences. Subjective tests have been conducted on usability and quality of the recommendation system. Luca Chiarandini, Massimiliano Zanoni, Augusto Sarti |
MMSP | 3 |
| 2010 | Geometric reconstruction of the environment from its response to multiple acoustic emissionsabstractIn this paper we propose a method for reconstructing the 2D geometry of the surrounding environment based on the signals acquired by a fixed microphone, when a series of acoustic stimula are produced in different positions in space. After estimating the Times Of Arrival (TOAs) of the reflective paths, we turn each TOA into a projective geometric constraint that can be used for determining the locations of the reflectors. The result consists of a collection of planar surfaces that correspond to the reflectors' locations. In this paper we present the whole processing chain and prove its effectiveness through experimental results. Fabio Antonacci, Augusto Sarti, Stefano Tubaro |
ICASSP | 2 |
| 2010 | Visibility-based beam tracing for soundfield renderingabstractIn this paper we present a visibility-based beam tracing solution for the simulation of the acoustics of environment that makes use of a projective geometry representation. More specifically, projective geometry turns out to be useful for the pre-computation of the visibility among all the reflectors in the environment. The simulation engine has a straightforward application in the rendering of the acoustics of virtual environments using loudspeaker arrays. More specifically, the acoustic wavefield is conceived as a superposition of acoustic beams, whose parameters (i.e. origin, orientation and aperture) are computed using the fast beam tracing methodology presented here. This information is processed by the rendering engine to compute spatial filters to be applied to the loudspeakers within the array. Simulative results show that an accurate simulation of the acoustic wavefield can be obtained using this approach. Dejan Markovic, Antonio Canclini, Fabio Antonacci, Augusto Sarti, Stefano Tubaro |
MMSP | 4 |
| 2010 | Geometric calibration of distributed microphone arrays from acoustic source correspondencesabstractThis paper proposes a method that solves the problem of geometric calibration of microphone arrays. We consider a distributed system, in which each array is controlled by separate acquisition devices that do not share a common synchronization clock. Given a set of probing sources, e.g. loudspeakers, each array computes an estimate of the source locations using a conventional TDOA-based algorithm. These observations are fused together by the proposed method, in order to estimate the position and pose of one array with respect to the other. Unlike previous approaches, we explicitly consider the anisotropic distribution of localization errors. As such, the proposed method is able to address the problem of geometric calibration when the probing sources are located both in the near- and far-field of the microphone arrays. Experimental results demonstrate that the improvement in terms of calibration accuracy with respect to state-of-the-art algorithms can be substantial, especially in the far-field. S. Daniele Valente, Marco Tagliasacchi, Fabio Antonacci, Paolo Bestagini, Augusto Sarti, Stefano Tubaro |
MMSP | 5 |
| 2010 | Virtual Analog Modeling in the Wave-Digital DomainabstractWe refer to ¿Virtual Analog¿ (VA) as a wide class of digital implementations that are modeled after nonlinear analog circuits for generating or processing musical sounds. The reference analog system is therefore typically represented by a set of blocks that are connected with each other through electrical ports, and usually exhibits a nonlinear behavior. It therefore seems quite natural to consider Nonlinear Wave Digital modeling as a solid approach for the rapid prototyping of such systems. In this paper, we discuss how nonlinear wave digital modeling can be fruitfully used for this purpose, with particular reference to special blocks and connectors that allow us to overcome the implementational difficulties and potential limitations of such solutions. In particular, we address some issues that are typical of VA and physical modeling, concerning how to accommodate special blocks into WD structures, how to enable the interaction between different WD structures, and how to accommodate structural and topological changes on the fly. Giovanni De Sanctis, Augusto Sarti |
IEEE Trans. Speech Audio Process. | 2 |
| 2009 | Musical audio semantic segmentation exploiting analysis of prominent spectral energy peaks and multi-feature refinementabstractIn this paper we present a novel hierarchical and scalable three-stage algorithm to effectively perform musical audio semantic segmentation. In the first stage, the energy spectrum of the entire audio track is analyzed to find significant energy textures that may characterize different semantic segments; in the second and third stages, tonal and timbric features are used to refine the segmentation by moving or deleting segment boundaries. Experimental results on a set of 58 songs show that our algorithm is able to attain good semantic segmentation just after the first step, with a precision of 64% and a recall of 96%. After second step the precision increases to 79%; the best precision result is obtained after the third step, where a value of 85% is reached. In this step the minimum average recall value of 92% is obtained. P. Romano, Giorgio Prandi, Augusto Sarti, Stefano Tubaro |
ICASSP | 3 |
| 2009 | Geometric and radiometric modeling of 3D scenesabstractModeling of 3D scenes is a hot topic in computer vision from more that thirty years, and probably its history is longer than a century considering also photogrammetry. In the recent years the rapid technological improvements that characterized the acquisition devices (photo-cameras, video-cameras, ..), illumination devices (lasers, structured light sources) and computational units allowed the application of 3D shape estimation methods, based on image analysis techniques, in a wide set of applications. Furthermore real-time 3D analysis is becoming a common tool in virtual and augmented reality contexts. Aim of this presentation is a rapid description of recent major advances on geometric and radiometric modeling of 3D scenes based on image analysis. Marco Marcon, Augusto Sarti, Stefano Tubaro |
ICME | 2 |
| 2009 | Geometric calibration of distributed microphone arraysabstractComputational auditory scene analysis exploits signals acquired by means of microphone arrays. In some circumstances, more than one array is deployed in the same environment. In order to effectively fuse the information gathered by each array, the relative location and pose of the arrays needs to be obtained solving a problem of geometric inter-array calibration. We consider the case where the arrays do not share a synchronous clock, which impairs the use of time-difference of arrival measures across arrays. Conversely, each array produces an acoustic image, which describes the energy of acoustic signals received from different directions. We jointly consider acoustic images acquired by the different arrays and adapt computer vision techniques to solve the calibration problem, thus estimating the location and pose of microphone arrays sensing the same auditory scene. We evaluate the robustness of the calibration process in a simulated environment and we investigate the effect of the various system parameters, namely the number of probing signal locations, the resolution of the acoustic images, the non-ideal intra-array calibration. Alessandro Redondi, Marco Tagliasacchi, Fabio Antonacci, Augusto Sarti |
MMSP | 4 |
| 2008 | Resource constrained efficient acoustic source localization and tracking using a distributed network of microphonesabstractIn this paper we present an efficient method to perform acoustic source localization and tracking using a distributed network of microphones. In this scenario, there is a trade-off between the localization performance and the expense of resources: in fact, a minimization of the localization error would require to use as many sensors as possible; at the same time, as the number of microphones increases, the cost of the network inevitably tends to grow, while in practical applications only a limited amount of resources is available. Therefore, at each time instant only a subset of the sensors should be enabled in order to meet the cost constraints. We propose a heuristic method for the optimal selection of this subset of microphones, using as distortion metrics the Cramer-Rao lower bound (CRLB) and as cost function the total distance between the selected sensors. The heuristic approach has been compared to an optimal algorithm, which searches the best sensor configuration among the full set of microphones, while satisfying the cost constraint. The proposed heuristic algorithm yields similar performance w.r.t. the full-search procedure, but at a much less computational cost. We show that this method can be used effectively in an acoustic source tracking application. Giuseppe Valenzise, Giorgio Prandi, Marco Tagliasacchi, Augusto Sarti |
ICASSP | 4 |
| 2008 | Uncalibrated view synthesis from Relative Affine Structure based on planes parallelismabstractThis paper focuses on the generation of physically valid views from two or more uncalibrated images acquired by standard cameras. The problem is faced without trying to yield a three dimensional reconstruction of the imaged scene, which would be unfeasible without the exact knowledge of the positions of the cameras in the Euclidean frame where the scene is to be described. Instead, starting from the previous works of Shashua and Navab on relative affine structure (1996) and the article of Fusiello on views synthesis from uncalibrated views (2007) we propose a novel approach that does not require the presence of a plane at infinity to define the homography between two views but merely the parallelism between couples of planes. This allows our approach to be applied to numerous scenes where two parallel planes can be defined (indoor scenes, straight streets and avenues). Experiments with synthetic images illustrate the approach. Stefano Tebaldini, Marco Marcon, Augusto Sarti, Stefano Tubaro |
ICIP | 3 |
| 2008 | Fast PDE approach to surface reconstruction from large cloud of points
Marco Marcon, Luca Piccarreta, Augusto Sarti, Stefano Tubaro |
Comput. Vis. Image Underst. | 3 |
| 2008 | 3D Motion from structures of points, lines and planes
Andrea Dell'Acqua, Augusto Sarti, Stefano Tubaro |
Image Vis. Comput. | 2 |
| 2008 | Fast Tracing of Acoustic Beams and Paths Through Visibility LookupabstractThe beam tracing method can be used for the fast tracing of a large number of acoustic paths through a direct lookup of a special tree-like data structure (beam tree) that describes the iterated visibility information from one specific position. This structure describes the branching of bundles of rays (beams) as they encounter reflectors in their paths. For this reason, beam tracing is suitable for real-time acoustic rendering even when the receiver is moving. In this paper, we propose a novel technique that enables the fast tracing of a large number of acoustic beams through the iterative lookup of a special data structure that describes the global visibility between reflectors. The method enables the immediate generation of the beam tree corresponding to an arbitrary source location, which can then be used for path tracing through direct lookup. In practice, this technique generalizes the traditional beam-tracing method as it makes it suitable for real-time acoustic rendering not just when the receiver is moving but also when the source is moving. The method enables real-time modeling of acoustic propagation and real-time auralization in complex 2-D and 2-Dtimes1-D environments (e.g., vertical walls limited by horizontal floor and ceiling), which makes it suitable for applications of real-time virtual acoustics, immersive gaming, and advanced acoustic rendering. Some experimental results show the effectiveness of fast beam tracing with respect to the state of the art in acoustic beam tracing. Fabio Antonacci, Marco Foco, Augusto Sarti, Stefano Tubaro |
IEEE Trans. Speech Audio Process. | 3 |
| 2007 | Tracking of two acoustic sources in reverberant environments using a particle swarm optimizerabstractIn this paper we consider the problem of tracking multiple acoustic sources in reverberant environments. The solution that we propose is based on the combination of two techniques. A blind source separation (BSS) method known as TRINICON [5] is applied to the signals acquired by the microphone arrays. The TRINICON de-mixing filters are used to obtain the Time Differences of Arrival (TDOAs), which are related to the source location through a nonlinear function. A particle filter is then applied in order to localize the sources. Particles move according to a swarm-like dynamics, which significatively reduces the number of particles involved with respect to traditional particle filter. We discuss results for the case of two sources and four microphone pairs. In addition, we propose a method, based on detecting source inactivity, which overcomes the ambiguities that intrinsically arise when only two microphone pairs are used. Experimental results demonstrate that the average localization error on a variety of pseudo-random trajectories is around 40 cm when the T60reverberation time is 0.6s. Fabio Antonacci, Davide Riva, Augusto Sarti, Marco Tagliasacchi, Stefano Tubaro |
AVSS | 3 |
| 2007 | Scream and gunshot detection and localization for audio-surveillance systemsabstractThis paper describes an audio-based video surveillance system which automatically detects anomalous audio events in a public square, such as screams or gunshots, and localizes the position of the acoustic source, in such a way that a video-camera is steered consequently. The system employs two parallel GMM classifiers for discriminating screams from noise and gunshots from noise, respectively. Each classifier is trained using different features, chosen from a set of both conventional and innovative audio features. The location of the acoustic source which has produced the sound event is estimated by computing the time difference of arrivals of the signal at a microphone array and using linear-correction least square localization algorithm. Experimental results show that our system can detect events with a precision of 93% at a false rejection rate of 5% when the SNR is 10dB, while the source direction can be estimated with a precision of one degree. A real-time implementation of the system is going to be installed in a public square of Milan. Giuseppe Valenzise, Luigi Gerosa, Marco Tagliasacchi, Fabio Antonacci, Augusto Sarti |
AVSS | 5 |
| 2006 | 3-D Body Posture Tracking For Human Action Template MatchingabstractIn this paper we present a novel approach to 3-D human action classification based on the analysis of volumetric data obtained form the joint processing of video sequences acquired by a multiple-camera system. The use of volumetric data makes the system very robust and avoids problems related the typical human body self-occlusions and motion ambiguities, very common in an independent camera-by-camera analysis. A shape descriptor of a human body is obtained in order to capture only posture-dependent characteristics and its outputs at each time instant are collected together in action feature matrices. The use of dynamic time warping approach for action template matching accounts for possible temporal nonlinear distortions among different instances of the same gesture and allows gesture classification Massimiliano Pierobon, Marco Marcon, Augusto Sarti, Stefano Tubaro |
ICASSP (2) | 3 |
| 2006 | A Robust Method for the Estimation of Reliable Wide Baseline CorrespondencesabstractIn this paper we present a complete method to retrieve reliable correspondences among wide baseline images, that is images of the same scene/object acquired from very different viewpoints. We propose a solution based on matching of affine co-variant features, composed by the following four steps: interest region detection, normalization, description and matching. In our method we implemented improved versions of some techniques recently introduced in the literature: the MSER detector (maximally stable extremal regions) and SIFT and RIFT descriptors (scale/rotation invariant feature transform). After a general introduction to the wide baseline problems and a summary of the recent state-of-the-art solutions, we illustrate the proposed method detailing the added improvements, then we present some experimental results obtained on wide baseline images. Francesco Colletto, Marco Marcon, Augusto Sarti, Stefano Tubaro |
ICIP | 3 |
| 2006 | On the Modeling of Motion in Wyner-Ziv Video CodingabstractIn the past few years, a number of practical video coding schemes following distributed source coding principles have emerged. One of the main goals of distributed video coding (DVC) is to enable a flexible distribution of the computational complexity between the encoder and the decoder, while approaching the coding efficiency of conventional closed-loop motion-compensated predictive codecs. In this paper we perform a rate-distortion analysis of a well-known Wyner-Ziv architecture, while focusing our attention on the impact of the motion modeling that is used for generating the side information at the decoder. Our analysis is structured according to a Kalman filtering problem and it allows us to compare three different scenarios: motion estimation at the encoder; motion interpolation at the decoder; and motion extrapolation at the decoder. Marco Tagliasacchi, Stefano Tubaro, Augusto Sarti |
ICIP | 3 |
| 2006 | Memory Extraction From Dynamic Scattering Junctions in Wave Digital StructuresabstractIn this letter, we show that a computable tree-like interconnection of parallel/series wave digital (WD) adaptors with memory (i.e., characterized by reflection filters instead of reflection coefficients) is equivalent to a like interconnection of standard (memoryless) adaptors whose peripheral ports are connected to mutators (two-port adaptors with memory). In proving all this, we provide a methodology for "extracting the memory" from a macro-adaptor, which can be fruitfully employed to simplify the implementation of WD structures Augusto Sarti, Giovanni De Sanctis |
IEEE Signal Process. Lett. | 1 |
| 2005 | Colored visual tags: a robust approach for augmented realityabstractThis paper presents a robust method for fast visual tags reading, suitable for augmented reality (AR) environments. Tag detection is based on well known tools of image-processing, but their combination, together with the use of colored markers, allows a robust recognition even with low-cost CMOS or CCD cameras and in poorly illuminated environments. In particular the color mix and the structure of the tag are quite unusual in common environments and can be easily detected with color filtering and geometric analysis. The proposed tag carries binary information encoded in its structure: in the presented implementation a 32-bit code with 12 parity bits is encoded in the tag but extensions to longer codes can be easily devised. Andrea Dell'Acqua, Marco Ferrari 0001, Marco Marcon, Augusto Sarti, Stefano Tubaro |
AVSS | 4 |
| 2005 | Clustering of human actions using invariant body shape descriptor and dynamic time warpingabstractWe propose a human action clustering method based on a 3D representation of the body in terms of volumetric coordinates. Features representing body postures are extracted directly from 3D data, making the system inherently insensitive to viewpoint dependence, motion ambiguities and self-occlusions. An invariant shape descriptor of human body is obtained in order to capture only posture-dependent characteristics, despite possible differences in translation, orientation, scale and body size. Frame-by-frame descriptions, generated from a gesture sequence, are collected together in matrices. Clustering of action matrices is eventually performed, and through a dynamic time warping (while computing the distance metric), we gain independence from possible temporal nonlinear distortions among different instances of the same gesture. Massimiliano Pierobon, Marco Marcon, Augusto Sarti, Stefano Tubaro |
AVSS | 3 |
| 2005 | Efficient source localization and tracking in reverberant environments using microphone arraysabstractIn this paper, we propose an algorithm for acoustic source localization and tracking that is suitable for reverberant environments. The approach that we propose is based on the iterative identification of the FIR channels that link source and microphones through an LMS method (multi-channel LMS), but we propose additional solutions that significantly improve this method in terms of computational efficiency and localization reliability, without affecting its convergence properties. This is achieved through a modified block-wise implementation of the approach combined with Kalman filtering. We also show the results of extensive comparative testing using novel performance parameters for the assessment of localization reliability. Fabio Antonacci, Davide Lonoce, Marco Motta, Augusto Sarti, Stefano Tubaro |
ICASSP (4) | 4 |
| 2005 | 3D object modeling with a voxelset carving approachabstractIn the past few years several systems for object reconstruction based on the analysis of 2D images have been proposed. In order for such systems to be of practical use, the 3D data extraction process is expected to be fast and reliable. In this paper we propose a general approach for the reconstruction of complete 3D objects based on a mesh fusion algorithm. Every surface patch is obtained as a depth map using an algorithm based on graph cuts theory. Each depth map is then triangulated before using it in a fusion algorithm based on a voxel-set carving approach. The result of the process is a closed mesh representing the object surface with sub-voxel resolution. Giovanni Dainese, Marco Marcon, Augusto Sarti, Stefano Tubaro |
ICIP (1) | 3 |
| 2005 | Combining MCTF with distributed source codingabstractMotion compensated temporal filtering (MCTF) has proved to be an efficient coding tool in the design of open-loop scalable video codecs. In this paper we propose a MCTF video coding scheme based on lifting where the prediction step is implemented using PRISM (power efficient, robust, high compression syndrome-based multimedia coding), a video coding framework built on distributed source coding principles. We study the effect of integrating the update step at the encoder or at the decoder side. We show that the latter approach allows improving the quality of the side information exploited during decoding. We present the analytical results obtained by modeling the video signal along the motion trajectories as an AR(1) process showing that the update step at the decoder allows to half the contribution of the quantization noise. We also include experimental results with real video data that demonstrate the potential of this approach when the video sequences are coded at low bitrates. Marco Tagliasacchi, Stefano Tubaro, Augusto Sarti |
ICIP (1) | 3 |
| 2004 | Action modeling with volumetric dataabstractIn this paper we propose and test an action recognition algorithm in which the images of the scene captured by a significant number of cameras are first used to generate a volumetric representation of a moving human body in terms of voxsets by means of volumetric intersection. The recognition stage is then performed directly on 3D data, allowing the system to avoid critical problems like viewpoint dependence and motion trajectory variability. Suitable features are extracted from the voxset representing the body and fed to a classical hidden Markov model to produce a finite-state description of the motion. Fabio Cuzzolin, Augusto Sarti, Stefano Tubaro |
ICIP | 2 |
| 2004 | Scalable coding of variable size blocks motion vectorsabstractIn this paper we discuss an algorithm that is able to provide a scalable (multiresolution) representation of the motion field information. It has been recently demonstrated that, for an open loop wavelet video coder, it is possible to use at the decoder side a scaled version of the original motion information and the residual coefficients computed with the full resolution/quality motion field without incurring into drift. We propose a fully scalable wavelet based video coder that performs motion estimation-compensation in the wavelet domain. In particular the coding scheme is specifically designed for variable size block matching algorithms. In this scenario motion vectors are distributed across an irregular lattice according to a quadtree structure. The developed system allows a scalable representation of the motion field and a flexible allocation of the bit budget between motion and residual data. The simulations that we have carried out show, at low bit-rates, a significant gain of the proposed approach with respect to the case in which the motion information is coded lossless. Davide Maestroni, Augusto Sarti, Marco Tagliasacchi, Stefano Tubaro |
ICIP | 2 |
| 2004 | Fast in-band motion estimation with variable size block matchingabstractIn this paper we propose a fast motion estimation technique that works in the wavelet domain. The computational cost of the algorithm turns out to be proportional to the linear size of the search window instead of its area. We complete our proposal with a variable-size block-matching scheme in the wavelet domain. We integrated the motion estimation algorithm in a fully-scalable wavelet in-band prediction coder inspired by the IB-MCTF (in-band motion compensation temporal filtering) proposed in J.C Ye et al., (2003). Comparative tests prove that our coder provides the same quality level than IB-MCTF at a very reduced computational cost. Moreover, although our method turns out to match the performance of MCTF-EZBC (P. Chen, 2003) in terms of PSNR, it clearly outperforms it in terms of perceptual quality, as it is completely free from blocking artifacts. Davide Maestroni, Augusto Sarti, Marco Tagliasacchi, Stefano Tubaro |
ICIP | 2 |
| 2004 | Area matching based on belief propagation with applications to face modeling
Davide Onofrio, Augusto Sarti, Stefano Tubaro |
ICIP | 2 |
| 2004 | Accurate and fast audio-realistic rendering of sounds in virtual environmentsabstractIn this paper we propose a novel method for real-time auralization of sounds in complex environments using visibility diagrams. The method accounts for both specular and diffracted reflections with receivers and sources that are free to move. Our solution concentrates in a pre-processing phase all the operations that can be conducted without knowledge of either source or receiver locations. In fact we pre-compute a set of visibility diagrams and diffracted beam trees. Once the source location is specified, we can determine the reflective beam trees through a simple lookup process on the visibility diagrams. The additional knowledge of the receiver location allows us to immediately generate all reflective and diffractive paths that link source and receiver. Fabio Antonacci, Marco Foco, Augusto Sarti, Stefano Tubaro |
MMSP | 3 |
| 2004 | Invariant action classification with volumetric dataabstractWe propose an action recognition algorithm in which the image sequences capturing a moving human body produced by a significant number of cameras are first used to generate a volumetric representation of the body by means of volumetric intersection. Classification is then performed directly on 3D data, making the system inherently insensitive to viewpoint dependence and motion trajectory variability. Suitable features are extracted from the voxset approximating the body, and fed to a hidden Markov model to produce a finite-state description of the motion. The Kullback-Leibler distance is finally used to classify new sequences. Fabio Cuzzolin, Augusto Sarti, Stefano Tubaro |
MMSP | 2 |
| 2004 | Detection of linear objects in GPR data
Andrea Dell'Acqua, Augusto Sarti, Stefano Tubaro, Luigi Zanzi |
Signal Process. | 2 |
| 2003 | Three-view camera calibration using geometric algebraabstractIn a former work of ours C. Defferara et al. (2002), we proposed a new way to express and interpret the epipolar constraint using Geometric Algebra, and we derived from it a novel and efficient 2-view camera calibration technique. In this paper we extend this GA approach to the 3-view case. After expressing the trifocal constraint in terms of bivectors and trivectors, we provide an alternative geometric interpretation of the coefficients of the trifocal tensor. On the basis of that, we propose a novel solution for the simultaneous determination of the focal lengths of the cameras and the rigid motion between three views. Andrea Dell'Acqua, Augusto Sarti, Stefano Tubaro |
ICIP (1) | 2 |
| 2002 | Robust real-time intrusion detection with fuzzy classificationabstractWe propose a novel system for indoor video surveillance. It is able to detect and track moving objects even in the presence of significant variations of scene illumination. After a preliminary analysis and clustering of temporal changes in the video sequence, the algorithm performs a classification based on fuzzy logic, aimed at identifying moving regions that really correspond to unexpected objects in the scene. The proposed approach tends to discard shadows, reflections and luminance profile changes due to illumination variations. One key feature of our system is its modest computation complexity, which allows it to operate in real-time on a common PC platform. The system has been tested on a wide variety of situations, proving its effectiveness and robustness. Giovanni Milanesi, Augusto Sarti, Stefano Tubaro |
ICIP (3) | 2 |
| 2002 | New perspectives on camera calibration using geometric algebraabstractWe propose a new approach to the camera self-calibration problem, based on geometric algebra. After an introduction on the adopted Clifford algebra framework, we provide new insight on the epipolar constraint as defined in terms of bivectors. On the basis of that, we propose a novel solution for the simultaneous determination of the focal lengths of the cameras and the rigid motion between views. Augusto Sarti, Claudio Defferara, Fabio Negroni, Stefano Tubaro |
ICIP (2) | 1 |
| 2002 | Image-based surface modeling: a multi-resolution approach
Augusto Sarti, Stefano Tubaro |
Signal Process. | 1 |
| 2002 | Detection and characterisation of planar fractures using a 3D Hough transform
Augusto Sarti, Stefano Tubaro |
Signal Process. | 1 |
| 2001 | Image-based object modeling: a multiresolution level-set approachabstractWe propose an image-based 3D modeling method based on a multi-resolution evolution of a level-set of a volumetric function, steered by the texture mismatch between views. A key feature of this approach is that it operates in an adaptive multi-resolution fashion, which boosts up the computational efficiency. A. Colosimo, Augusto Sarti, Stefano Tubaro |
ICIP (2) | 2 |
| 2001 | A multiresolution level-set approach to surface fusionabstractWe propose a novel solution to the problem of combining a number of partial reconstructions through a process of 3D "patchworking". Our method is able to seamlessly "sew" the surface overlaps together, and to reasonably "mend" all the holes that remain after surface assembly, which usually correspond to the non-visible portions of the object surface. The approach is based on the temporal evolution of the zero level set volumetric function, driven by surface curvature and distance from data and works entirely in a multi-resolution fashion. Augusto Sarti, Stefano Tubaro |
ICIP (2) | 1 |
| 2000 | Multi-Resolution Corner DetectionabstractIn this paper we propose a novel technique for wavelet-based multi-resolution corner detection. The method is based on a search of the zero crossings of the Laplacian at different scales along the line that describes the trajectory of the maxima as the scale varies. The proposed techniques allow us to achieve sub-pixel accuracy and provide us with useful scale information on the detected features. Federico Pedersini, Elena Pozzoli, Augusto Sarti, Stefano Tubaro |
ICIP | 3 |
| 2000 | Multi-Resolution Area MatchingabstractWe present a general and robust approach to the problem of close-range partial 3D reconstruction of objects from multi-resolution texture matching. The method is based on the progressive refinement of a parametric surface, which is described using an increasing number of radial functions. Federico Pedersini, Augusto Sarti, Stefano Tubaro |
ICIP | 2 |
| 2000 | Multicamera motion estimation for high-accuracy 3D reconstruction
Federico Pedersini, Pasquale Pigazzini, Augusto Sarti, Stefano Tubaro |
Signal Process. | 3 |
| 2000 | Visible surface reconstruction with accurate localization of object boundariesabstractA common limitation of many techniques for 3-D reconstruction from multiple perspective views is the poor quality of the results near the object boundaries. The interpolation process applied to "unstructured" 3-D data ("clouds" of non-connected 3-D points) plays a crucial role in the global quality of the 3-D reconstruction. We present a method for interpolating unstructured 3-D data, which is able to perform a segmentation of such data into different data sets that correspond to different objects. The algorithm is also able to perform an accurate localization of the boundaries of the objects. The method is based on an iterative optimization algorithm. As a first step, a set of surfaces and boundary curves are generated for the various objects. Then, the edges of the original images are used for refining such boundaries as best as possible. Experimental results with real data are presented for proving the effectiveness of the proposed algorithm. Federico Pedersini, Augusto Sarti, Stefano Tubaro |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 1999 | Estimation of Radiometric Parameters for a Realistic Rendering of 3D ModelsabstractIn order to obtain realistic results in the rendering of a 3D model, we need accurate information on both structure (shape) and radiometric properties (reflectivity) of its surfaces. In this paper, we approach the problem of correctly estimating the radiometric characteristics of an imaged surface through the analysis of several of its views. The images are acquired with one or more CCD cameras that move around the object, and the 3D model is assumed available (it can be constructed from the available images). The adopted reflectivity model is non-Lambertian as it takes into account both the diffuse lobe and the specular lobe. The method implements a robust technique which is able to estimate the reflectivity parameters of the object's surfaces even in the presence of modeling imperfections and/or when the relative camera-object position and orientation are only approximately known. Federico Pedersini, Luca Piccarreta, Augusto Sarti, Stefano Tubaro |
ICIP (4) | 3 |
| 1999 | Accurate and simple geometric calibration of multi-camera systems
Federico Pedersini, Augusto Sarti, Stefano Tubaro |
Signal Process. | 2 |
| 1998 | Accurate Feature Detection and Matching for the Tracking of Calibration Parameters in Multi-Camera Acquisition SystemsabstractThe 3D reconstruction's quality of multiple-camera acquisition systems is strongly influenced by the accuracy of the camera calibration procedure. The acquisition of long sequences is, in fact, very sensitive to mechanical shocks, vibrations and thermal changes on cameras and supports, as they could result in a significant drift of the camera parameters. In this paper we propose a technique which is able to keep track of the camera parameters and, whenever possible, to correct them accordingly. This technique does not need any a-priori knowledge or test objects to be placed in the scene, but exploits features that are already present in the scene itself. In fact it performs an accurate detection, matching and back-projection of luminance corners and spots in the scene space. Experimental results on real sequences are reported in order to prove the ability of the proposed technique to detect a change in the calibration and to re-calibrate the camera setup with an accuracy that depends on the number of available feature points. Federico Pedersini, Augusto Sarti, Stefano Tubaro |
ICIP (2) | 2 |
| 1998 | Combined Surface Interpolation and Object Segmentation for Automatic 3-D Scene ReconstructionabstractA common limitation of many techniques for 3D reconstruction from multiple perspective views is the poor quality of the results near the object boundaries. The interpolation process applied to "unstructured" 3D data ("clouds" of non-connected 3D points) plays a crucial role in the global quality of the 3D reconstruction. We present a method for interpolating unstructured 3D data, which is able to perform a segmentation of such data into different data sets that correspond to different objects. The algorithm is also able to perform an accurate localization of the boundaries of the objects. The method is based on an iterative optimization algorithm. As a first step, a set of surfaces and boundary curves are generated for the various objects. Then, the edges of the original images are used for refining such boundaries as best as possible. Experimental results with real data are presented for proving the effectiveness of the proposed algorithm. Federico Pedersini, Augusto Sarti, Stefano Tubaro |
ICIP (2) | 2 |
| 1998 | 3D area matching with arbitrary multiview geometry
Federico Pedersini, Pasquale Pigazzini, Augusto Sarti, Stefano Tubaro |
Signal Process. Image Commun. | 3 |
| 1998 | Improving the performance of edge localization techniques through error compensation
Federico Pedersini, Augusto Sarti, Stefano Tubaro |
Signal Process. Image Commun. | 2 |
| 1997 | Egomotion Estimation of a Multicamera System Through Line CorrespondenceabstractIn this paper we propose a method for estimating the egomotion of a calibrated multi-camera system from an analysis of the luminance edges. The method works entirely in the 3D space as all edges of each one set of views are previously localized, matched and back-projected onto the object space. In fact, it searches for the rigid motion that best merges the sets of 3D contours extracted from each one of the multi-views. The method uses both straight and curved 3D contours. Federico Pedersini, Augusto Sarti, Stefano Tubaro |
ICIP (2) | 2 |
| 1997 | Robust Area MatchingabstractWe propose a general and robust solution to the problem of close-range 3D reconstruction of objects from stereo correspondence of luminance profiles. The method does not require a particular camera geometry, and can be implemented with an arbitrary number of CCD cameras. Its robustness can be mainly attributed to the physicality of the matching process, which is performed in the 3D space, while taking both geometric and radiometric distortions into account. Extensive tests have been performed over a variety of real scenes, using a calibrated trinocular camera system. Federico Pedersini, Augusto Sarti, Stefano Tubaro |
ICIP (2) | 2 |
| 1997 | Estimation and Compensation of Subpixel Edge Localization ErrorabstractWe propose and analyze a method for improving the performance of subpixel edge localization (EL) techniques through compensation of the systematic portion of the localization error. The method is based on the estimation of the EL characteristic through statistical analysis of a test image and is independent of the EL technique in use. Federico Pedersini, Augusto Sarti, Stefano Tubaro |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1997 | Joint automatic design of prefilter and decimation grid for the reduction of spectral redundancy in 2D digital signals
Federico Pedersini, Augusto Sarti, Stefano Tubaro |
Signal Process. | 2 |
| 1996 | Combined motion and edge analysis for a layer-based representation of image sequencesabstractWe propose a method for the motion-congruent segmentation of image sequences based on both motion field and luminance information. In order to do so, the affine motion models are determined by analyzing the motion field through a clustering procedure, while their regions of validity are determined by an MRF-based region estimator. This last block performs a pixel-wise re-assignment of a limited number of affine models (those determined via clustering). Federico Pedersini, Augusto Sarti, Stefano Tubaro |
ICIP (1) | 2 |
| 1996 | 3D motion estimation of a trinocular system for a full-3D object reconstructionabstractMotion estimation of a calibrated multi-ocular acquisition system can be performed either on two-dimensional data, by applying a rigidity constraint to a set of matched points on monocular views, or on three-dimensional data, by determining the rigid motion that best overlaps homologous 3D points obtained through stereo matching and back-projection. We propose, analyze and compare these two possible solutions, and present a low-cost high-accuracy full-3D reconstruction system based on multiple trinocular views at standard TV resolution. Federico Pedersini, Augusto Sarti, Stefano Tubaro |
ICIP (2) | 2 |
| 1995 | Synthesis of virtual views using non-Lambertian reflectivity models and stereo matchingabstractA technique for the synthesis of virtual views of a 3D scene, starting from images taken by a calibrated multicamera system, is proposed and tested. Surface interpolation is performed over a set of 3D edges, computed with stereometric algorithms, and additional curvature-tuning points, scattered where the reflectivity model is sufficiently reliable. The 3D coordinates of the tuning points are computed by minimizing the MSE between the available real views and the corresponding synthetic views. Synthesis is finally carried out by reprojecting on the new image plane the estimated object surface over which texture-mapping of the reflection-corrected luminance function has been performed. Texture correction, which uses an estimate of a non-Lambertian reflectivity model, is done in such a way to simulate the migration of reflections due to the change of viewpoint. The technique has been tested on real images, producing realistic synthesized views. Federico Pedersini, Augusto Sarti, Stefano Tubaro |
ICIP | 2 |
| 1994 | Nonlinearity compensation in digital radio systemsabstractDigital radio links with bandlimited pulses exhibit a severe performance degradation when the transmitter high power amplifier operates near saturation. To cope with the increase of nonlinear intersymbol interference due to the amplifier nonlinearities, a discrete-time Volterra system can be used to process the transmitted data. We present an efficient technique for implementing adaptive data predistorters with memory based on a discrete-time Volterra system composed of digital linear filters and memoryless nonlinear devices working at the symbol rate. Third- and fifth-order structures are proposed and a system performance evaluation is presented for several realistic situations.> Giovanni Lazzarin, Silvano Pupolin, Augusto Sarti |
IEEE Trans. Commun. | 3 |