EDBT 2026 Demo / reviewers in the wild / expert
Archontis Politis
dblp:136/5019
· DBLP profile ↗
28ranked-venue papers
4as first author
16since 2021 · last 2025
0000-0002-0595-2356ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 18 · 3 first-author · 9 since 2021Artificial intelligence and machine learning · 12 · 1 first-author · 10 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Gen-A: Generalizing Ambisonics Neural Encoding to Unseen Microphone ArraysabstractUsing deep neural networks (DNNs) for encoding of microphone array (MA) signals to the Ambisonics spatial audio format can surpass certain limitations of established conventional methods, but existing DNN-based methods need to be trained separately for each MA. This paper proposes a DNN-based method for Ambisonics encoding that can generalize to arbitrary MA geometries unseen during training. The method takes as inputs the MA geometry and MA signals and uses a multi-level encoder consisting of separate paths for geometry and signal data, where geometry features inform the signal encoder at each level. The method is validated in simulated anechoic and reverberant conditions with one and two sources. The results indicate improvement over conventional encoding across the whole frequency range for dry scenes, while for reverberant scenes the improvement is frequency-dependent. Mikko Heikkinen, Archontis Politis, Konstantinos Drossos, Tuomas Virtanen |
ICASSP | 2 |
| 2025 | Inter-Speaker Relative Cues for Text-Guided Target Speech Extraction
Wang Dai, Archontis Politis, Tuomas Virtanen |
INTERSPEECH | 2 |
| 2025 | Attractor-Based Speech Separation of Multiple Utterances by Unknown Number of Speakers
Archontis Politis, Konstantinos Drossos, Tuomas Virtanen |
INTERSPEECH | 2 |
| 2024 | Neural Ambisonics Encoding For Compact Irregular Microphone ArraysabstractAmbisonics encoding of microphone array signals can enable various spatial audio applications, such as virtual reality or telepresence, but it is typically designed for uniformly-spaced spherical microphone arrays. This paper proposes a method for Ambisonics encoding that uses a deep neural network (DNN) to estimate a signal transform from microphone inputs to Ambisonics signals. The approach uses a DNN consisting of a U-Net structure with a learnable preprocessing as well as a loss function consisting of mean average error, spatial correlation, and energy preservation components. The method is validated on two microphone arrays with regular and irregular shapes having four microphones, on simulated reverberant scenes with multiple sources. The results of the validation show that the proposed method can meet or exceed the performance of a conventional signal-independent Ambisonics encoder on a number of error metrics. Mikko Heikkinen, Archontis Politis, Tuomas Virtanen |
ICASSP | 2 |
| 2024 | Perceptually-Motivated Spatial Audio Codec for Higher-Order Ambisonics CompressionabstractScene-based spatial audio formats, such as Ambisonics, are playback system agnostic and may therefore be favoured for delivering immersive audio experiences to a wide range of (potentially unknown) devices. The number of channels required to deliver high spatial resolution Ambisonic audio, however, can be prohibitive for low-bandwidth applications. Therefore, this paper proposes a compression codec, which is based upon the parametric higher-order Directional Audio Coding (HO-DirAC) model. The encoder downmixes the higher-order Ambisonic (HOA) input audio into a reduced number of signals, which are accompanied by perceptually-motivated scene parameters. The downmixed audio is coded using a perceptual audio coder, whereas the parameters are grouped into perceptual bands, quantized, and downsampled. On the decoder side, low Ambisonic orders are fully recovered. Not fully recoverable HOA components are synthesized according to the parameters. The results of a listening test indicate that the proposed parametric spatial audio codec can improve the adopted perceptual audio coder, especially at low to medium-high bitrates, when applied to fifth-order HOA signals. Christoph Hold, Leo McCormack, Archontis Politis, Ville Pulkki |
ICASSP | 3 |
| 2024 | Attention-Driven Multichannel Speech Enhancement in Moving Sound Source ScenariosabstractCurrent multichannel speech enhancement algorithms typically assume a stationary sound source, a common mismatch with reality that limits their performance in real-world scenarios. This paper focuses on attention-driven spatial filtering techniques designed for dynamic settings. Specifically, we study the application of linear and nonlinear attention-based methods for estimating time-varying spatial covariance matrices used to design the filters. We also investigate the direct estimation of spatial filters by attention-based methods without explicitly estimating spatial statistics. The clean speech clips from WSJ0 are employed for simulating speech signals of moving speakers in a reverberant environment. The experimental dataset is built by mixing the simulated speech signals with multichannel real noise from CHiME-3. Evaluation results show that the attention-driven approaches are robust and consistently outperform conventional spatial filtering approaches in both static and dynamic sound environments. Archontis Politis, Tuomas Virtanen |
ICASSP | 2 |
| 2024 | Dynamic Processing Neural Network Architecture for Hearing Loss CompensationabstractThis paper proposes neural networks for compensating sensorineural hearing loss. The aim of the hearing loss compensation task is to transform a speech signal to increase speech intelligibility after further processing by a person with a hearing impairment, which is modeled by a hearing loss model. We propose an interpretable model called dynamic processing network, which has a structure similar to band-wise dynamic compressor. The network is differentiable, and therefore allows to learn its parameters to maximize speech intelligibility. More generic models based on convolutional layers were tested as well. The performance of the tested architectures was assessed using spectro-temporal objective index (STOI) with hearing-threshold noise and hearing aid speech intelligibility (HASPI) metrics. The dynamic processing network gave a significant improvement of STOI and HASPI in comparison to popular compressive gain prescription rule Camfit. A large enough convolutional network could outperform the interpretable model with the cost of larger computational load. Finally, a combination of the dynamic processing network with convolutional neural network gave the best results in terms of STOI and HASPI. Szymon Drgas, Lars Bramslow, Archontis Politis, Gaurav Naithani, Tuomas Virtanen |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2024 | Compression of Higher-Order Ambisonic Signals Using Directional Audio CodingabstractDelivering high-quality spatial audio in the Ambisonics format requires extensive data bandwidth, which may render it inaccessible for many low-bandwidth applications. Existing widely-available multi-channel audio compression codecs are not designed to consider the characteristic inter-channel relations inherent to the Ambisonics format, and thus may not leverage this knowledge to optimise the compression. Therefore, this article proposes a spatial audio compression algorithm, based on a novel reformulation of the Higher-Order Directional Audio Coding (HO-DirAC) method, which is specifically intended for compressing higher-order Ambisonic audio streams. The methodology builds upon the concept of a spherical filter bank acting in the spherical harmonic domain. This results in directionally constrained sound-field estimates and parameterization, which may be utilized to reconstruct the input Ambisonic signals with minimal perceived loss of quality. The results of a listening experiment indicate high perceptual quality when using six or more audio transport channels to deliver fifth-order (36 channels) Ambisonic sound scenes. The proposed formulation is also designed with low computational complexity in mind and may therefore be well suited for compressing Ambisonic sound scenes for a wide range of applications. Christoph Hold, Ville Pulkki, Archontis Politis, Leo McCormack |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2024 | Binaural Sound Source Distance Estimation and Localization for a Moving ListenerabstractIn this paper, we investigate the tasks of binaural source distance estimation (SDE) and direction-of-arrival estimation (DOAE) using motion-based cues in a scenario with a walking listener. On top of performing both tasks as separate problems, we study two methods of solving the joint task of simultaneous source distance estimation and localization (SDEL), with a single model. Experiments are conducted for three different scenarios: a static receiver; a static receiver with a rotating head; and a freely moving listener inside a room. The study proposes rotation and translation features to include information about the receiver's motion during model training and studies the effects of these on the final performance. The work includes extended simulation of three datasets containing numerous testing scenarios for sound sources, covering a wide range of DOAs and a source-to-receiver distance up to 15 m. Results are further analyzed with respect to room reverberation, walking speed, as well as source-to-receiver distance. The presented outcomes show large improvements in both DOA and distance estimation for a model that uses motion-based cues as compared with a static scenario. These include a decrease of 9.50° in DOA and 1.56m in distance errors for a joint model, followed by 16.17° and 0.17m for separate models. Daniel Krause 0001, Guillermo García-Barrios, Archontis Politis, Annamaria Mesaros |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2024 | Speaker Distance Estimation in Enclosures From Single-Channel AudioabstractDistance estimation from audio plays a crucial role in various applications, such as acoustic scene analysis, sound source localization, and room modeling. Most studies predominantly center on employing a classification approach, where distances are discretized into distinct categories, enabling smoother model training and achieving higher accuracy but imposing restrictions on the precision of the obtained sound source position. Towards this direction, in this paper we propose a novel approach for continuous distance estimation from audio signals using a convolutional recurrent neural network with an attention module. The attention mechanism enables the model to focus on relevant temporal and spectral features, enhancing its ability to capture fine-grained distance-related information. To evaluate the effectiveness of our proposed method, we conduct extensive experiments using audio recordings in controlled environments with three levels of realism (synthetic room impulse response, measured response with convolved speech, and real recordings) on four datasets (our synthetic dataset, QMULTIMIT, VoiceHome-2, and STARSS23). Experimental results show that the model achieves an absolute error of 0.11 meters in a noiseless synthetic scenario. Moreover, the results showed an absolute error of about 1.30 meters in the hybrid scenario. The algorithm's performance in the real scenario, where unpredictable environmental factors and noise are prevalent, yields an absolute error of approximately 0.50 meters. For reproducible research purposes we make model, code, and synthetic datasets available at https://github.com/michaelneri/audio-distance-estimation Michael Neri, Archontis Politis, Daniel Krause 0001, Marco Carli, Tuomas Virtanen |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2023 | How to (Virtually) Train Your Speaker LocalizerabstractLearning-based methods have become ubiquitous in speaker localization. Existing systems rely on simulated training sets for the lack of sufficiently large, diverse and annotated real datasets. Most room acoustics simulators used for this purpose rely on the image source method (ISM) because of its computational efficiency. This paper argues that carefully extending the ISM to incorporate more realistic surface, source and microphone responses into training sets can significantly boost the real-world performance of speaker localization systems. It is shown that increasing the training-set realism of a state-of-the-art direction-of-arrival estimator yields consistent improvements across three different real test sets featuring human speakers in a variety of rooms and various microphone arrays. An ablation study further reveals that every added layer of realism contributes positively to these improvements. Prerak Srivastava, Antoine Deleforge, Archontis Politis, Emmanuel Vincent 0001 |
INTERSPEECH | 3 |
| 2023 | STARSS23: An Audio-Visual Dataset of Spatial Recordings of Real Scenes with Spatiotemporal Annotations of Sound EventsabstractWhile direction of arrival (DOA) of sound events is generally estimated from multichannel audio data recorded in a microphone array, sound events usually derive from visually perceptible source objects, e.g., sounds of footsteps come from the feet of a walker. This paper proposes an audio-visual sound event localization and detection (SELD) task, which uses multichannel audio and video information to estimate the temporal activation and DOA of target sound events. Audio-visual SELD systems can detect and localize sound events using signals from a microphone array and audio-visual correspondence. We also introduce an audio-visual dataset, Sony-TAu Realistic Spatial Soundscapes 2023 (STARSS23), which consists of multichannel audio data recorded with a microphone array, video data, and spatiotemporal annotation of sound events. Sound scenes in STARSS23 are recorded with instructions, which guide recording participants to ensure adequate activity and occurrences of sound events. STARSS23 also serves human-annotated temporal activation labels and human-confirmed DOA labels, which are based on tracking results of a motion capture system. Our benchmark results demonstrate the benefits of using visual object positions in audio-visual SELD tasks. The data is available at https://zenodo.org/record/7880637. Kazuki Shimada, Archontis Politis, Parthasaarathy Sudarsanam, Daniel Krause 0001, Kengo Uchida, Sharath Adavanne, Aapo Hakala, Yuichiro Koyama, Naoya Takahashi, Shusuke Takahashi, Tuomas Virtanen, Yuki Mitsufuji |
NeurIPS | 2 |
| 2022 | Parametric Ambisonic Encoding of Arbitrary Microphone ArraysabstractThis article proposes a parametric signal-dependent method for the task of encoding microphone array signals into Ambisonic signals. The proposed method is presented and evaluated in the context of encoding a simulated seven-sensor microphone array, which is mounted on an augmented reality headset device. Given the inherent flexibility of the Ambisonics format, and its popularity within the context of such devices, this array configuration represents a potential future use case for Ambisonic recording. However, due to its irregular geometry and non-uniform sensor placement, conventional signal-independent Ambisonic encoding is particularly limited. The primary aims of the proposed method are to obtain Ambisonic signals over a wider frequency band-width, and at a higher spatial resolution, than would otherwise be possible through conventional signal-independent encoding. The proposed method is based on a multi-source sound-field model and employs spatial filtering to divide the captured sound-field into its individual source and directional ambient components, which are subsequently encoded into the Ambisonics format at an arbitrary order. It is demonstrated through both objective and perceptual evaluations that the proposed parametric method outperforms conventional signal-independent encoding in the majority of cases. Leo McCormack, Archontis Politis, Raimundo Gonzalez, Tapio Lokki, Ville Pulkki |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2021 | Spherical Decomposition of Arbitrary Scattering Geometries for Virtual Acoustic EnvironmentsabstractA method is proposed to encode the acoustic scattering of objects for virtual acoustic applications through a multiple-input and multiple-output framework. The scattering is encoded as a matrix in the spherical harmonic domain, and can be re-used and manipulated (rotated, scaled and translated) to synthesize various sound scenes. The proposed method is applied and validated using Boundary Element Method simulations which shows accurate results between references and synthesis. The method is compatible with existing frameworks such as Ambisonics and image source methods. Raimundo Gonzalez, Archontis Politis, Tapio Lokki |
DAFx | 2 |
| 2021 | Parametric Spatial Audio Effects Based on the Multi-Directional Decomposition of Ambisonic Sound ScenesabstractDecomposing a sound-field into its individual components and respective parameters can represent a convenient first-step towards offering the user an intuitive means of controlling spatial audio effects and sound-field modification tools. The majority of such tools available today, however, are instead limited to linear combinations of signals or employ a basic single-source parametric model. Therefore, the purpose of this paper is to present a parametric framework, which seeks to overcome these limitations by first dividing the sound-field into its multi-source and ambient components based on estimated spatial parameters. It is then demonstrated that by manipulating the spatial parameters prior to reproducing the scene, a number of sound-field modification and spatial audio effects may be realised; including: directional warping, listener translation, sound source tracking, spatial editing workflows and spatial side-chaining. Many of the effects described have also been implemented as real-time audio plug-ins, in order to demonstrate how a user may interact with such tools in practice. Leo McCormack, Archontis Politis, Ville Pulkki |
DAFx | 2 |
| 2021 | Overview and Evaluation of Sound Event Localization and Detection in DCASE 2019abstractSound event localization and detection is a novel area of research that emerged from the combined interest of analyzing the acoustic scene in terms of the spatial and temporal activity of sounds of interest. This paper presents an overview of the first international evaluation on sound event localization and detection, organized as a task of the DCASE 2019 Challenge. A large-scale realistic dataset of spatialized sound events was generated for the challenge, to be used for training of learning-based approaches, and for evaluation of the submissions in an unlabeled subset. The overview presents in detail how the systems were evaluated and ranked and the characteristics of the best-performing systems. Common strategies in terms of input features, model architectures, training approaches, exploitation of prior knowledge, and data augmentation are discussed. Since ranking in the challenge was based on individually evaluating localization and event classification performance, part of the overview focuses on presenting metrics for the joint measurement of the two, together with a reevaluation of submissions using these new metrics. The new analysis reveals submissions that performed better on the joint task of detecting the correct type of event close to its original location than some of the submissions that were ranked higher in the challenge. Consequently, ranking of submissions which performed strongly when evaluated separately on detection or localization, but not jointly on both, was affected negatively. Archontis Politis, Annamaria Mesaros, Sharath Adavanne, Toni Heittola, Tuomas Virtanen |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2020 | Frequency-Dependent Directional Feedback Delay NetworkabstractA recent publication introduced the Directional Feedback Delay Network, a parametric artificial reverberation algorithm capable of producing direction-dependent energy decay. This method extends the capabilities of Feedback Delay Networks by using multichannel delay-line groups and a spatial transform to produce direction-dependent reverberation in the Ambisonics domain. In this paper, we present a modified formulation of the Directional Feedback Delay Network method that allows both frequency-and direction-dependent reverberation. Multichannel delay-line groups are used to manipulate signals incident on a spherical grid through independent recursive signal paths, while an early reflection module gives a physically-motivated spatial distribution of input signals in the system. The overall number of delay lines is reduced from the previous formulation, and the design allows flexibility to favor a lower computational cost over accuracy. Benoit Alary, Archontis Politis |
ICASSP | 2 |
| 2020 | Multichannel Singing Voice Separation by Deep Neural Network Informed DOA Constrained CMNMFabstractThis work addresses the problem of multichannel source separation combining two powerful approaches, multichannel spectral factorization with recent monophonic deep learning (DL) based spectrum inference. Individual source spectra at different channels are estimated with a Masker-Denoiser twin network, able to model long-term temporal patterns of a musical piece. The monophonic source spectrograms are used within a spatial covariance mixing model based on complex-valued multichannel non-negative matrix factorization (CMNMF) that predicts the spatial characteristics of each source. The proposed framework is evaluated on the task of singing voice separation with a large multichannel dataset. Experimental results show that our joint DL+CMNMF method outperforms both the individual monophonic DL-based separation and the multichannel CMNMF baseline methods. Antonio Jesús Muñoz-Montoro, Archontis Politis, Konstantinos Drossos, Julio J. Carabias-Orti |
MMSP | 2 |
| 2020 | Blind reverberation time estimation from ambisonic recordingsabstractReverberation time is an important room acoustic parameter, useful for many acoustic signal processing applications. Most of the existing work on blind reverberation time estimation focuses on the single-channel case. However, the recent developments and interest on immersive audio have brought to the market a number of spherical microphone arrays, together with the usage of ambisonics as a standard spatial audio convention. This work presents a novel blind reverberation time estimation method, which specifically targets ambisonic recordings, a field that remained unexplored to the best of our knowledge. Experimental validation on a synthetic reverberant dataset shows that the proposed algorithm outperforms state-of-the-art methods under most evaluation criteria in low noise conditions. Andrés Pérez-López, Archontis Politis, Emilia Gómez |
MMSP | 2 |
| 2019 | Sharpening of Angular Spectra Based on a Directional Re-assignment Approach for Ambisonic Sound-field VisualisationabstractA method for computing and sharpening angular spectra, derived from low-order ambisonic signals, is presented in this paper, which is intended for high-resolution directional sound-field visualisation. The method relies on a re-assignment principle, whereby the directional energy for each grid point is assigned to a new direction, which corresponds to a direction-of-arrival (DoA) estimate within a spatially-localised region, centred around the respective grid point. This leads to the concentration of energy around the true sources, and hence, to sharper angular spectra than that of steered response power (SRP) beamformers of maximum directivity, with the same order of ambisonic input. It is demonstrated that the proposed method, when using low-order input, can achieve similar results to the SRP approach of much higher order. Leo McCormack, Archontis Politis, Ville Pulkki |
ICASSP | 2 |
| 2019 | Local Time-Domain Spherical Harmonic Spatial Encoding for Wave-Based Acoustic SimulationabstractVolumetric time-domain simulation methods, such as the finite difference time domain method, allow for a fine-grained representation of the dynamics of the acoustic field. A key feature of such methods is complete access to the computed field, normally represented over a Cartesian grid. Simple solutions to the problem of extracting spatially encoded signals, necessary in virtual acoustics applications, result. In this letter, a simple time-domain representation of spatially encoded spherical harmonic signals is written directly in terms of spatial derivatives of the acoustic field at the receiver location. In a discrete setting, encoded signals may be obtained, at very low computational cost and latency, using local approximations with minimal number of grid points, and avoiding large convolutions and frequency-domain block processing of previous approaches. Numerical results illustrating receiver directivity and computed time-domain responses are presented, as well as numerical solution drift associated with repeated time integration. Stefan Bilbao, Archontis Politis, Brian Hamilton |
IEEE Signal Process. Lett. | 2 |
| 2018 | COMPASS: Coding and Multidirectional Parameterization of Ambisonic Sound ScenesabstractCurrent methods for immersive playback of spatial sound content aim at flexibility in terms of encoding and decoding, abstracting the two from the recording or playback setup. Ambisonics constitutes such a method, that is however signal-independent, and at low spatial resolutions fails to provide appropriate spatialization cues to the listener, with potential severe colouration effects and localization ambiguity. We present a new signal-dependent method for parametric analysis and synthesis of ambisonic sound scenes that takes advantage of the flexibility of Ambisonics as a spatial audio format, while improving reproduction. The proposed approach considers a more general acoustic model than previous proposals, with multiple source signals and a non isotropic ambient component. According to a listening test using headphones, the method is perceived closer to binaural reference sound scenes than ambisonic playback. Archontis Politis, Sakari Tervo, Ville Pulkki |
ICASSP | 1 |
| 2018 | Multichannel Sound Event Detection Using 3D Convolutional Neural Networks for Learning Inter-channel FeaturesabstractIn this paper, we propose a stacked convolutional and recurrent neural network (CRNN) with a 3D convolutional neural network (CNN) in the first layer for the multichannel sound event detection (SED) task. The 3D CNN enables the network to simultaneously learn the inter-and intra-channel features from the input multichannel audio. In order to evaluate the proposed method, multichannel audio datasets with different number of overlapping sound sources are synthesized. Each of this dataset has a four-channel first-order Ambisonic, binaural, and single-channel versions, on which the performance of SED using the proposed method are compared to study the potential of SED using multichannel audio. A similar study is also done with the binaural and single-channel versions of the real-life recording TUT-SED 2017 development dataset. The proposed method learns to recognize overlapping sound events from multichannel features faster and performs better SED with a fewer number of training epochs. The results show that on using multichannel Ambisonic audio in place of single-channel audio we improve the overall F-score by 7.5%, overall error rate by 10% and recognize 15.6% more sound events in time frames with four overlapping sound sources. Sharath Adavanne, Archontis Politis, Tuomas Virtanen |
IJCNN | 2 |
| 2016 | Applications of 3D spherical transforms to personalization of head-related transfer functionsabstractHead-related transfer functions (HRTFs) depend on the shape of the human head and ears, motivating HRTF personalization methods that detect and exploit morphological similarities between subjects in an HRTF database and a new user. Prior work determined similarity from sets of morphological parameters. Here we propose a non-parametric morphological similarity based on a harmonic expansion of head scans. Two 3D spherical transforms are explored for this task, and an appropriate shape similarity metric is defined. A case study focusing on personalisation of interaural time differences (ITDs) is conducted by applying this similarity metric on a database of 3D head scans. Archontis Politis, Mark R. P. Thomas, Hannes Gamper, Ivan Tashev |
ICASSP | 1 |
| 2015 | Direction-of-arrival and diffuseness estimation above spatial aliasing for symmetrical directional microphone arraysabstractA method for direction-of-arrival (DOA) and diffuseness estimation is presented, which proves to be effective above the spatial aliasing frequency of the microphone array in use. The method assumes symmetrical circular or spherical arrays of directional microphones or microphone mounted on a rigid baffle, and it exploits the inherent directionality of the array at high frequencies. The DOA and diffuseness estimators are shown to exhibit low estimation error above the spatial aliasing limit, compared to commonly used intensity-based estimators. A low-error broadband scheme that combines the intensity-based method and the new one is proposed for the ranges below and above aliasing respectively. Archontis Politis, Symeon Delikaris-Manias, Ville Pulkki |
ICASSP | 1 |
| 2015 | ATSI: Augmented and Tangible Sonic InteractionabstractThis paper presents ATSI, a system for sonic augmentation of physical objects with spatialized sounds and their control by gestural interaction. The implementation combines tracking of the users' hands and head with commodity hardware, and binaural spatialized sound rendering over headphones. The user may attach sounds to common objects and the system maintains the correct spatial auditory perspective inside the augmented scene during the actions of attaching sounds, picking and moving the sounding objects, or exploring the scene. Attention has been given to the binaural rendering of distance cues to support the above actions with perceptual realism in small scale environments and for many objects. Through a user study assessing the localization accuracy that can be achieved with the system, we show that sound rendering performance looks appropriate for applications such as auditory displays for context and object specific information, sonic design for architectural planning and interior design, and music applications. Roberto Pugliese, Archontis Politis, Tapio Takala |
TEI | 2 |
| 2015 | Direction of Arrival Estimation of Reflections from Room Impulse Responses Using a Spherical Microphone ArrayabstractThis paper studies the direction of arrival estimation of reflections in short time windows of room impulse responses measured with a spherical microphone array. Spectral-based methods, such as multiple signal classification (MUSIC) and beamforming, are commonly used in the analysis of spatial room impulse responses. However, the room acoustic reflections are highly correlated or even coherent in a single analysis window and this imposes limitations on the use of spectral-based methods. Here, we apply maximum likelihood (ML) methods, which are suitable for direction of arrival estimation of coherent reflections. These methods have been earlier developed in the linear space domain and here we present the ML methods in the context of spherical microphone array processing and room impulse responses. Experiments are conducted with simulated and real data using the em32 Eigenmike. The results show that direction estimation with ML methods is more robust against noise and less biased than MUSIC or beamforming. Sakari Tervo, Archontis Politis |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2013 | A directional diffuse reverberation model for excavated tunnels in rockabstractAcoustic impulse responses of an excavated tunnel were measured. Analysis of the impulse responses shows that they are very diffuse from the start. A reverberator suitable for reproducing this type of response is proposed. The input signal is first comb-filtered and then convolved with a sparse noise sequence of the same length as the filter's delay line. An IIR loop filter inside the comb filter determines the decay rate of the response and is derived from the Yule-Walker approximation of the measured frequency-dependent reverberation time. The particular sparse noise sequence proposed in this work combines three velvet noise sequences, two of which have time-varying weights. To simulate the directional soundfield in a tunnel, the use of multiple such reverberators, each associated with a virtual source distributed evenly around the listener, is suggested. The proposed tunnel acoustics simulation can be employed in gaming, in film sound, or in working machine simulators. Sami Oksanen, Julian Parker, Archontis Politis, Vesa Välimäki |
ICASSP | 3 |