EDBT 2026 Demo / reviewers in the wild / expert
Konrad Kowalczyk
dblp:78/8062
· DBLP profile ↗
40ranked-venue papers
6as first author
24since 2021 · last 2025
0000-0002-7834-6920ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 32 · 3 first-author · 21 since 2021Artificial intelligence and machine learning · 19 · 3 first-author · 12 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Investigation of Whisper ASR Hallucinations Induced by Non-Speech AudioabstractHallucinations of deep neural models are amongst key challenges in automatic speech recognition (ASR). In this paper, we investigate hallucinations of the Whisper ASR model induced by non-speech audio segments present during inference. By inducting hallucinations with various types of sounds, we show that there exists a set of hallucinations that appear frequently. We then study hallucinations caused by the augmentation of speech with such sounds. Finally, we describe the creation of a bag of hallucinations (BoH) that allows to remove the effect of hallucinations through the post-processing of text transcriptions. The results of our experiments show that such post-processing is capable of reducing word error rate (WER) and acts as a good safeguard against problematic hallucinations. Mateusz Baranski, Jan Jasinski, Julitta Bartolewska, Stanislaw Kacprzak, Marcin Witkowski, Konrad Kowalczyk |
ICASSP | 6 |
| 2025 | Clustering-based Hard Negative Sampling for Supervised Contrastive Speaker Verification
Piotr Masztalski, Michal Romaniuk, Jakub Zak, Mateusz Matuszewski, Konrad Kowalczyk |
INTERSPEECH | 5 |
| 2025 | Joint Diarization and Separation Using SepFormer With Non-Autoregressive Attractors
Magdalena Rybicka, Konrad Kowalczyk, Thomas Thebaud, Najim Dehak, Jesús Villalba 0001 |
IEEE Signal Process. Lett. | 2 |
| 2024 | Heightceleb - An Enrichment of Voxceleb Dataset With Speaker Height InformationabstractPrediction of speaker’s height is of interest for voice forensics, surveillance, and automatic speaker profiling. Until now, TIMIT has been the most popular dataset for training and evaluation of the height estimation methods. In this paper, we introduce HeightCeleb, an extension to VoxCeleb, which is the dataset commonly used in speaker recognition tasks. This enrichment consists in adding information about the height of all 1251 speakers from VoxCeleb that has been extracted with an automated method from publicly available sources. Such annotated data will enable the research community to utilize freely available speaker embedding extractors, pre-trained on VoxCeleb, to build more efficient speaker height estimators. In this work, we describe the creation of the HeightCeleb dataset and show that using it enables to achieve state-of-the-art results on the TIMIT test set by using simple statistical regression methods and embeddings obtained with a popular speaker model (without any additional fine-tuning). Stanislaw Kacprzak, Konrad Kowalczyk |
SLT | 2 |
| 2024 | Reverberant Source Separation Using NTF With Delayed Subsources and Spatial PriorsabstractSpeech signals recorded by distant microphones are often contaminated with room reverberation and signals of interfering speakers. This article addresses the problem of joint source separation and dereverberation using multichannel nonnegative tensor factorization (NTF) in which late reverberant components are modeled using the so-called delayed subsources. The article formulates two distinct signal models of the time-frequency spectrum of the multichannel microphone mixture, in which reverberation is modeled either independently for each source using delayed source variances or jointly using delayed microphone signals. In addition, it defines computationally efficient variants of these two methods with a simplified spatial model in which spatial properties of the late reverberant components are estimated jointly for all delays. For each of the four distinct algorithms, the article first formulates a maximum a posteriori (MaP) estimator based on the NTF model with the localization prior over the mixing matrix that is suitable for the estimation of the early reverberation (primarily the direct-path) signals in a reverberant environment. Next it derives update equations for the four resulting expectation-maximization algorithms, which are thoroughly evaluated and shown to outperform similar state-of-the-art approaches. The results of experimental evaluations, performed using real and simulated data, for determined, over-determined and under-determined scenarios, indicate superior performance of the proposed processing over state-of-the-art in terms of standard source separation and dereverberation metrics. Mieszko Fras, Konrad Kowalczyk |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2024 | On Ambisonic Source Separation With Spatially Informed Non-Negative Tensor FactorizationabstractThis article presents a Non-negative Tensor Factorization based method for sound source separation from Ambisonic microphone signals. The proposed method enables the use of prior knowledge about the Directions-of-Arrival (DOAs) of the sources, incorporated through a constraint on the Spatial Covariance Matrix (SCM) within a Maximum a Posteriori (MAP) framework. Specifically, this article presents a detailed derivation of four algorithms that are based on two types of cost functions, namely the squared Euclidean distance and the Itakura-Saito divergence, which are then combined with two prior probability distributions on the SCM, that is the Wishart and the Inverse Wishart. The experimental evaluation of the baseline Maximum Likelihood (ML) and the proposed MAP methods is primarily based on first-order Ambisonic recordings, using four different source signal datasets, three with musical pieces and one containing speech utterances. We consider underdetermined, determined, as well as over-determined scenarios by separating two, four and six sound sources, respectively. Furthermore, we evaluate the proposed algorithms for different spherical harmonic orders and at different reverberation time levels, as well as in non-ideal prior knowledge conditions, for increasingly more corrupted DOAs. Overall, in comparison with beamforming and a state-of-the-art separation technique, as well as the baseline ML methods, the proposed MAP approach offers superior separation performance in a variety of scenarios, as shown by the analysis of the experimental evaluation results, in terms of the standard objective separation measures, such as the SDR, ISR, SIR and SAR. Mateusz Guzik, Konrad Kowalczyk |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2024 | End-to-End Neural Speaker Diarization With Non-Autoregressive AttractorsabstractDespite many recent developments in speaker diarization, it remains a challenge and an active area of research to make diarization robust and effective in real-life scenarios. Well-established clustering-based methods are showing good performance and qualities. However, such systems are built of several independent, separately optimized modules, which may cause non-optimum performance. End-to-end neural speaker diarization (EEND) systems are considered the next stepping stone in pursuing high-performance diarization. Nevertheless, this approach also suffers limitations, such as dealing with long recordings and scenarios with a large (more than four) or unknown number of speakers in the recording. The appearance of EEND with encoder-decoder-based attractors (EEND-EDA) enabled us to deal with recordings that contain a flexible number of speakers thanks to an LSTM-based EDA module. A competitive alternative over the referenced EEND-EDA baseline is the EEND with non-autoregressive attractor (EEND-NAA) estimation, proposed recently by the authors of this article. NAA back-end incorporates k-means clustering as part of the attractor estimation and an attractor refinement module based on a Transformer decoder. However, in our previous work on EEND-NAA, we assumed a known number of speakers, and the experimental evaluation was limited to 2-speaker recordings only. In this article, we describe in detail our recent EEND-NAA approach and propose further improvements to the EEND-NAA architecture, introducing three novel variants of the NAA back-end, which can handle recordings containing speech of a variable and unknown number of speakers. Conducted experiments include simulated mixtures generated using the Switchboard and NIST SRE datasets and real-life recordings from the CALLHOME and DIHARD II datasets. In experimental evaluation, the proposed systems achieve up to 51% relative improvement for the simulated scenario and up to 15% for real recordings over the baseline EEND-EDA. Magdalena Rybicka, Jesús Villalba 0001, Thomas Thebaud, Najim Dehak, Konrad Kowalczyk |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2023 | Convolutive NTF for Ambisonic Source Separation under Reverberant ConditionsabstractThis paper presents a Non-negative Tensor Factorization (NTF) based sound source separation method with a novel convolutive Spatial Covariance Matrix (SCM) model, that is suitable for use with reverberant Ambisonic signals. The presented solution builds upon a previous work on SHD SCM-based NTF, but unlike the original, non-convolutive approach, it avoids the problem encountered when the analysis window is too short to capture the dominant part of the reverberant signal. Here we introduce a novel convolutive SCM model that accounts for reverberation which spans over multiple time frames and then we derive the corresponding parameter update equations. In particular, this work considers several variants of these updates, describes the underlying motivation for each algorithm design choice and indicates the update rules, which offer the highest gain in Signal-to-Distortion Ratio (SDR). The proposed solution is evaluated against the original approach for various reverberation time values, number of sources and types of source signals, using simulated first-order Ambisonic recordings. The results of this preliminary study clearly indicate that the proposed method enables higher quality of separation compared with the reference, non-convolutive algorithm. Mateusz Guzik, Konrad Kowalczyk |
ICASSP | 2 |
| 2023 | Causal Signal-Based DCCRN with Overlapped-Frame Prediction for Online Speech EnhancementabstractThe aim of speech enhancement is to improve speech signal quality and intelligibility from a noisy microphone signal. In many applications, it is crucial to enable processing with small computational complexity and minimal requirements regarding access to future signal samples (look-ahead). This paper presents signal-based causal DCCRN that improves online single-channel speech enhancement by reducing the required look-ahead and the number of network parameters. The proposed modifications include complex filtering of the signal, application of overlapped-frame prediction, causal convolutions and deconvolutions, and modification of the loss function. Results of performed experiments indicate that the proposed model with overlapped signal prediction and additional adjustments, achieves similar or better performance than the original DCCRN in terms of various speech enhancement metrics, while it reduces the latency and network parameter number by around 30%. Julitta Bartolewska, Stanislaw Kacprzak, Konrad Kowalczyk |
INTERSPEECH | 3 |
| 2023 | Joint Blind Source Separation and Dereverberation for Automatic Speech Recognition using Delayed-Subsource MNMF with Localization Prior
Mieszko Fras, Marcin Witkowski, Konrad Kowalczyk |
INTERSPEECH | 3 |
| 2022 | Convolutional Weighted Minimum Mean Square Error Filter for Joint Source Separation and DereverberationabstractPractical scenarios with multiple simultaneously active speakers recorded using one or more microphones in reverberant rooms pose a challenging problem when the extraction of the desired speaker signal is sought for. The majority of techniques found in the literature facilitate either source separation or dereverberation, which can at best be performed as subsequent, cascade processing. Recently, a solution to the joint task has been proposed, which is known as the weighted power minimization distortionless response (WPD) beamformer. In this paper, we derive a convolutional multichannel filter which performs jointly optimum dereverberation and desired source signal extraction. We formulate a single optimization criterion which minimizes the convolutional source-variance weighted mean square error (CW-MMSE), thereby effectively unifying the weighted prediction error (WPE) based dereverberation and MMSE filtering for the desired source extraction from reverberant mixtures of speakers. Experimental results show a significant performance improvement over the compared state-of-the-art methods such as WPD for datasets with simulated and recorded impulse responses. Mieszko Fras, Marcin Witkowski, Konrad Kowalczyk |
ICASSP | 3 |
| 2022 | Wishart Localization Prior On Spatial Covariance Matrix In Ambisonic Source Separation Using Non-Negative Tensor FactorizationabstractThis paper presents an extension of the existing Non-negative Tensor Factorization (NTF) based method for sound source separation under reverberant conditions, formulated for Ambisonic microphone mixture signals. In particular, we address the problem of optimal exploitation of the prior knowledge concerning the source localization, through the formulation of a suitable Maximum a Posteriori (MAP) framework. Within the presented approach, the magnitude spectrograms are modelled by the NTF and the individual source Spatial Covariance Matrices (SCM) are approximated as a sum of anechoic Spherical Harmonic (SH) components, weighted with the so-called spatial selector. We constrain the SCM using the Wishart distribution, which leads to a new posterior probability and in turn to the derivation of the extended update rules. The proposed solution avoids the issues encountered in the original method, related to the empirical binary initialization strategy for the spatial selector weights, which due to multiplicative update rules may result in sound coming from certain directions not being taken into account. The proposed method is evaluated against the original algorithm and another recently proposed Expectation Maximization (EM) algorithm that also incorporates a spatial localization prior, showing improved separation performance in experiments with first-order Ambisonic recordings of musical instruments and speech utterances. Mateusz Guzik, Konrad Kowalczyk |
ICASSP | 2 |
| 2022 | Spoken Language Recognition with Cluster-Based ModelingabstractIn this study, we analyze the incorporation of cluster-based modeling into the language recognition systems, in which a single utterance is represented as an embedding, deploying widely used i-vectors and x-vectors. We compare the results obtained with a Cosine Distance Scoring, Gaussian Mixture Model, Logistic Regression, and the Mixture of von Misses-Fisher distributions with the classifiers based on the proposed approach which incorporates cluster-based sub-models. Experimental evaluation is performed on the i-vector embeddings from the NIST 2015 language recognition i-vector machine learning challenge and the x-vector embeddings from the Oriental Language Recognition 2020 Challenge (AP20-OLR). The experimental results clearly show that the proposed approach combined with discriminatively trained Logistic Regression classifier achieves notable improvements over the baseline systems, i.e., those without language sub-models, and that our approach is competitive to other systems reported in the literature. Stanislaw Kacprzak, Magdalena Rybicka, Konrad Kowalczyk |
ICASSP | 3 |
| 2022 | Refining DNN-based Mask Estimation using CGMM-based EM Algorithm for Multi-channel Noise ReductionabstractIn this paper, we present a method that allows to further improve speech enhancement obtained with recently introduced Deep Neural Network (DNN) models. We propose a multi-channel refinement method of time-frequency masks obtained with single-channel DNNs, which consists of an iterative Complex Gaussian Mixture Model (CGMM) based algorithm, followed by optimum spatial filtration. We validate our approach on time-frequency masks estimated with three recent deep learning models, namely DCUnet, DCCRN, and FullSubNet. We show that our method with the proposed mask refinement procedure allows to improve the accuracy of estimated masks, in terms of the Area Under the ROC Curve (AUC) measure, and as a consequence the overall speech quality of the enhanced speech signal, as measured by PESQ improvement, and that the improvement is consistent across all three DNN models. Julitta Bartolewska, Stanislaw Kacprzak, Konrad Kowalczyk |
INTERSPEECH | 3 |
| 2022 | Convolutive Weighted Multichannel Wiener Filter Front-end for Distant Automatic Speech Recognition in Reverberant Multispeaker Scenarios
Mieszko Fras, Marcin Witkowski, Konrad Kowalczyk |
INTERSPEECH | 3 |
| 2022 | NTF of Spectral and Spatial Features for Tracking and Separation of Moving Sound Sources in Spherical Harmonic Domain
Mateusz Guzik, Konrad Kowalczyk |
INTERSPEECH | 2 |
| 2022 | End-to-End Neural Speaker Diarization with an Iterative Refinement of Non-Autoregressive Attention-based Attractors
Magdalena Rybicka, Jesús Villalba 0001, Najim Dehak, Konrad Kowalczyk |
INTERSPEECH | 4 |
| 2022 | Convolutional Weighted Parametric Multichannel Wiener Filter for Reverberant Source SeparationabstractIn this letter, we address the problem of simultaneous separation and dereverberation of overlapped speech recorded in reverberant conditions. The majority of state-of-the-art techniques are tailored for either of the two problems, which for the joint task leads to sub-optimum performance or solutions which involve subsequent, cascade processing. In contrast, we propose a jointly optimum approach in which we formulate a single optimization criterion that minimizes variance of undesired signal components at the output of a convolutional filter, weighted with the desired speech variance, subject to a constraint which allows to control the amount of distortions in the estimated speech signal. We then derive a closed-form solution of the proposed convolutional weighted parametric multichannel Wiener (CW-PMW) filter which integrates linear-prediction based dereverberation and speech-distortion weighted Wiener filtering in a jointly optimum manner. The results of experiments performed using measured and simulated data indicate superior performance of the proposed approach in comparison with state-of-the-art, which includes sub-optimum cascades of optimum filters for individual tasks, as well as the recently presented jointly optimum, weighted power minimization distortionless response (WPD) beamformer. Mieszko Fras, Konrad Kowalczyk |
IEEE Signal Process. Lett. | 2 |
| 2021 | Maximum a Posteriori Estimator for Convolutive Sound Source Separation with Sub-Source Based NTF Model and the Localization Probabilistic Prior on the Mixing MatrixabstractIn this paper we present a method for the separation of sound source signals recorded using multiple microphones in a reverberant room. In particular, we propose a maximum a posteriori (MAP) estimator based on the multichannel nonnegative tensor factorization (NTF) model with the localization prior distribution on the mixing matrix, in which the latent data consists of the so-called sub-sources for an improved performance in a reverberant environment. For the proposed MAP estimator, we derive the sub-source based expectation maximization (EM) algorithm with the multiplicative update rules (MU) and the localization prior distribution (LP) on the mixing matrix (SSEM-MU-LP). We then perform several experiments for speech and instrumental sound sources recorded using two microphones, in determined and under-determined scenarios, and with different types of initialization of the model parameters. The results of these experiments clearly indicate a significant improvement of the proposed algorithm with the localization prior over the state-of-the-art NTF-based source separation algorithms, which can reach up to 50% in the signal-to-distortion ratio. Mieszko Fras, Konrad Kowalczyk |
ICASSP | 2 |
| 2021 | Combating Reverberation in NTF-Based Speech Separation Using a Sub-Source Weighted Multichannel Wiener Filter and Linear Prediction
Mieszko Fras, Marcin Witkowski, Konrad Kowalczyk |
Interspeech | 3 |
| 2021 | Spine2Net: SpineNet with Res2Net and Time-Squeeze-and-Excitation Blocks for Speaker Recognition
Magdalena Rybicka, Jesús Villalba 0001, Piotr Zelasko, Najim Dehak, Konrad Kowalczyk |
Interspeech | 5 |
| 2021 | Frame-based Maximum a Posteriori Estimation of Second-Order Statistics for Multichannel Speech Enhancement in Presence of NoiseabstractIn this paper, we present a multichannel noise reduction scheme which uses the minimum variance distortionless response (MVDR) beamformer based on the second-order statistics (SOS) of the source and noise signals estimated in a frame-wise fashion. The time-frequency masks, required for SOS estimation, correspond to the probabilities of speech presence and speech absence in the noisy observations, and they are found by the proposed frame-based maximum a posteriori (MAP) estimator with an inverse Wishart prior distribution. The derived expectation-maximization (EM) algorithm estimates the parameters of the assumed complex Gaussian mixture model (CGMM). The proposed approach is compared with an existing block-based method. Following the outline of mathematical differences between both processing schemes, we perform experimental evaluation. The obtained results indicate that the proposed frame-based approach outperforms block-based method by enabling stronger reduction of undesired noise, which in turn leads to better quality of the enhanced speech signal. Julitta Bartolewska, Konrad Kowalczyk |
MMSP | 2 |
| 2021 | Incorporation of Localization Information for Sound Source Separation in Spherical Harmonic DomainabstractThis paper concerns the problem of convolutive sound source separation from mutlichannel recordings made with a spherical microphone array. In particular, we formulate two state-of-the-art separation techniques based on Expectation Maximization (EM) and Nonnegative Tensor Factorization (NTF) in the spherical harmonic domain (SHD). Furthermore, we adjust and incorporate the Gaussian Localization Prior (GLP) to the proposed algorithms, which yields two variants of the derived methods. For the source signal reconstruction, a Minimum Variance Distortionless Response (MVDR) beamformer with a single-channel Wiener post-filter is employed. The performance comparison is based on experimental evaluation using micro-phone signals simulated with the image-source method in several scenarios, including diverse geometrical setup, different number of sources and various types of source signals, namely the recordings of speech utterances and musical instruments. The experimental results for the first-order ambisonic signals show that the proposed methods enable high-quality sound source separation in the spherical harmonic domain. In particular, we show that incorporation of the Gaussian Localization Prior to the proposed algorithms leads to a substantial improvement in separation performance. Mateusz Guzik, Mieszko Fras, Konrad Kowalczyk |
MMSP | 3 |
| 2021 | Split Bregman Approach to Linear Prediction Based Dereverberation With Enforced Speech SparsityabstractThe recordings of speech in enclosures are corrupted by reverberation caused by multipath wave propagation from the speaker to the distant microphones. In this letter, we address the problem of reducing the late part of room reverberation. The presented blind dereverberation method consists in multichannel linear prediction (MCLP) and enforces sparsity of the dereverberated speech by adopting the split Bregman approach. The proposed algorithm alternately solves two optimization problems, where the former cost function is derived by assuming that speech is modelled using a sparse prior distribution, while the latter optimization emphasizes speech sparsity by an additional incorporation of a weightedl1-norm of the output signal to the standard linear prediction based cost function. The results of experiments performed using simulated and measured room impulse responses for various reverberation time values indicate superior performance of the proposed sparse split Bregman (SSB) method over state-of-the-art non-sparse and sparse MCLP-based dereverberation methods in terms of standard evaluation measures and as pre-processsing to the automatic speech recognition. Marcin Witkowski, Konrad Kowalczyk |
IEEE Signal Process. Lett. | 2 |
| 2020 | Exploiting Rays in Blind Localization of Distributed Sensor ArraysabstractMany signal processing algorithms for distributed sensors are capable of improving their performance if the positions of sensors are known. In this paper, we focus on estimators for inferring the relative geometry of distributed arrays and sources, i.e. the setup geometry up to a scaling factor. Firstly, we present the Maximum Likelihood estimator derived under the assumption that the Direction of Arrival measurements follow the von Mises-Fisher distribution. Secondly, using unified notation, we show the relations between the cost functions of a number of state-of-the-art relative geometry estimators. Thirdly, we derive a novel estimator that exploits the concept of rays between the arrays and source event positions. Finally, we show the evaluation results for the presented estimators in various conditions, which indicate that major improvements in the probability of convergence to the optimum solution over the existing approaches can be achieved by using the proposed ray-based estimator. Szymon Wozniak, Konrad Kowalczyk |
ICASSP | 2 |
| 2020 | On Parameter Adaptation in Softmax-Based Cross-Entropy Loss for Improved Convergence Speed and Accuracy in DNN-Based Speaker Recognition
Magdalena Rybicka, Konrad Kowalczyk |
INTERSPEECH | 2 |
| 2019 | Passive Joint Localization and Synchronization of Distributed Microphone ArraysabstractCollaborative processing of signals recorded by distributed sensors brings about new opportunities as well as challenges to classical array signal processing. In this letter, we present a method for passive self-calibration of a distributed system in which each distributed node consists of an array of sensors. The proposed method estimates the positions and orientations of distributed sensor arrays, the synchronization timeline offsets as well as the positions of spatially distributed events emitted by an uncontrolled source or sources. The proposed two-step optimization consists in finding the maximum likelihood estimate of the relative geometry of a distributed system based on Directions of Arrival (DoAs) observed independently at each array. Next, the final positions of distributed arrays, the positions of acoustic events, and the synchronization offsets between the nodes are computed by solving a linear least square problem based on the observed Time Differences of Arrival (TDoAs) between the arrays and the relative positions estimated in the first optimization step. The evaluation results indicate that passive self-calibration of distributed sensor arrays can be conveniently achieved by observing acoustic events generated by an uncontrolled source based on the measured DoAs and TDoAs. Szymon Wozniak, Konrad Kowalczyk |
IEEE Signal Process. Lett. | 2 |
| 2017 | Audio Replay Attack Detection Using High-Frequency Features
Marcin Witkowski, Stanislaw Kacprzak, Piotr Zelasko, Konrad Kowalczyk, Jakub Galka |
INTERSPEECH | 4 |
| 2015 | Residual noise control using a parametric multichannel Wiener filterabstractMultichannel noise reduction techniques are commonly used in speech communication applications. In these applications, it is often desired to maintain a residual amount of background noise to avoid perceptually unpleasant artifacts, such as musical tones or time periods of complete silence. Noise reduction can be achieved by the parametric multichannel Wiener filter (PMWF), which provides a trade-off between speech distortion and noise reduction. To additionally control the maximum noise reduction, the PMWF can be decomposed into a spatial filter and a spectral gain, which is limited to a desired minimum value. Such decomposition is however only possible if the desired source power spectral density matrix is rank-one, which in general does not even hold for a single source in reverberant environments. In the proposed approach, we define the desired signal as a sum of the speech signal plus the desired residual noise, and derive an optimum filter in the minimum mean-square error sense. The resulting filter has the advantage that it enables direct control of the maximum noise reduction without the need for a gain limiting step and is furthermore applicable to desired signals of higher rank. We analyze the derived filter thoroughly and show its relation to the standard PMWF that results as a special case. Furthermore, we propose a solution for keeping the residual noise level constant in slowly time-varying noise fields. Sebastian Braun, Konrad Kowalczyk, Emanuël A. P. Habets |
ICASSP | 2 |
| 2015 | Binaural Reproduction of Finite Difference Simulations Using Spherical Array ProcessingabstractDue to its efficiency and simplicity, the finite-difference time-domain method is becoming a popular choice for solving wideband, transient problems in various fields of acoustics. So far, the issue of extracting a binaural response from finite difference simulations has only been discussed in the context of embedding a listener geometry in the grid. In this paper, we propose and study a method for binaural response rendering based on a spatial decomposition of the sound field. The finite difference grid is locally sampled using a volumetric array of receivers, from which a plane wave density function is computed and integrated with free-field head related transfer functions, in the spherical harmonics domain. The volumetric array is studied in terms of numerical robustness and spatial aliasing. Analytic formulas that predict the performance of the array are developed, facilitating spatial resolution analysis and numerical binaural response analysis for a number of finite difference schemes. Particular emphasis is placed on the effects of numerical dispersion on array processing and on the resulting binaural responses. Our method is compared to a binaural simulation based on the image method. Results indicate good spatial and temporal agreement between the two methods. Jonathan Sheaffer, Maarten van Walstijn, Boaz Rafaely, Konrad Kowalczyk |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2014 | Extended Kalman filter with probabilistic data association for multiple non-concurrent speaker localization in reverberant environmentsabstractAcoustic source localization and tracking (ASLT) in reverberant environments is a challenging task due to the multi-path propagation of acoustic waves. ASLT is often based on the use of a Kalman filter or a particle filter, with time-difference-of-arrival (TDOA) estimates used as measurements. In this work, we aim to track non-concurrent speakers by applying an extended Kalman filter (EKF) with probabilistic data association (PDA) that takes into account multiple measurements simultaneously. By using PDA, the inaccuracy of the measurements caused by room reflections and noise is explicitly taken into account. Unlike in typical approaches where the measurements consist of broadband TDOA estimates, the measurements in the proposed approach consist of multiple narrowband direction-of-arrival (DOA) estimates obtained from distributed microphone arrays. Experimental results demonstrate that incorporating PDA and using properly selected narrowband DOA estimates leads to a better tracking performance, as compared to the standard EKF with a single narrowband or broadband measurement. Soumitro Chakrabarty, Konrad Kowalczyk, Maja Taseska, Emanuël A. P. Habets |
ICASSP | 2 |
| 2014 | Estimation of Acoustic Reflection Coefficients Through Pseudospectrum MatchingabstractEstimating the geometric and reflective properties of the environment is important for a wide range of applications of space-time audio processing, from acoustic scene analysis to room equalization and spatial audio rendering. In this manuscript, we propose a methodology for frequency-subband in-situ estimation of the reflection coefficients of planar surfaces. This is a rather challenging task, as the reflection coefficients depend on the frequency and the angle of incidence and their estimate is highly sensitive to background noise and interfering sources. Our method is based on the assumption that we know the geometry of the reflectors; the position and the radiation pattern of the source; the position and the spatial response of the array. Applying beamforming algorithms on a single set of measured sensor data, we estimate the angular distribution of the acoustic energy (angular pseudospectrum) that impinges on a microphone array. We then apply a two-step iterative estimation technique based on an Expectation-Maximization (EM) algorithm. The first step estimates the scaling factors. The second one infers the reflection coefficients from the scaling factors. Under the assumption of additive white Gaussian noise, we finally determine the reflection coefficients with a Maximum Likelihood (ML) estimation method. The effectiveness and the accuracy of the proposed technique are assessed through experiments based on measured data. Dejan Markovic, Konrad Kowalczyk, Fabio Antonacci, Christian Hofmann 0001, Augusto Sarti, Walter Kellermann |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2013 | Blind System Identification Using Sparse Learning for TDOA Estimation of Room ReflectionsabstractLocalization of early room reflections can be achieved by estimating the time-differences-of-arrival (TDOAs) of reflected waves between elements of a microphone array. For an unknown source, we propose to apply sparse blind system identification (BSI) methods to identify the acoustic impulse responses, from which the TDOAs of temporally sparse reflections are estimated. The proposed time- and frequency-domain adaptive algorithms based on crossrelation formulation are regularized by incorporating an l1-norm sparseness constraint, which is realized using a split Bregman method. These algorithms are shown to outperform standard crossrelation-based BSI techniques when estimating TDOAs of reflections in the presence of background noise. Konrad Kowalczyk, Emanuël A. P. Habets, Walter Kellermann, Patrick A. Naylor |
IEEE Signal Process. Lett. | 1 |
| 2011 | On 2D localization of reflectors using robust beamforming techniquesabstractThis paper presents a method for the localization of reflectors in an acoustic environment, using robust beamforming techniques and a cylindrical microphone array, for which an intuitive and highly efficient three-step procedure is proposed. First, the directions of ar rival (DOAs) corresponding to the sound source and reflectors are estimated by a robust Minimum Variance Distortionless Response (MVDR) beamformer. Next, signals originating from the estimated DOAs are extracted by a robust superdirective beamformer, from which time differences of arrival (TDOAs) of major reflections are estimated by crosscorrelation analysis. Finally, by using additional information about the position of the direct sound source relative to the array, the positions of the reflective boundaries of the room can be inferred. Experiments based on real measurements carried out in a moderately reverberant room show the effectiveness of this method. Edwin Mabande, Haohai Sun, Konrad Kowalczyk, Walter Kellermann |
ICASSP | 3 |
| 2011 | Joint DOA and TDOA estimation for 3D localization of reflective surfaces using eigenbeam MVDR and spherical microphone arraysabstractMethods of 3D direction of arrival (DOA) estimation, coherent source detection and reflective surface localization are studied, based on recordings by a spherical microphone array. First, the spherical harmonics domain minimum variance distortionless response (EB-MVDR) beamformer is employed for the localization of broadband coherent sources, which is characterized by simpler frequency focusing matrices than the corresponding element-space implementation, and by a higher resolution than conventional spherical array beamformers. After the DOA estimation step, the source signals are extracted by EB-MVDRs. Then, by computing the crosscorrelation functions between the extracted signals, the coherent sources are detected and their time differences of arrival (TDOA) are estimated. Given the positions of the array and the reference source, and the estimated DOA and TDOA of the coherent sources, the positions of the major reflectors can be inferred. Experimental results in a real room validate the proposed method. Haohai Sun, Edwin Mabande, Konrad Kowalczyk, Walter Kellermann |
ICASSP | 3 |
| 2011 | Room Acoustics Simulation Using 3-D Compact Explicit FDTD SchemesabstractThis paper presents methods for simulating room acoustics using the finite-difference time-domain (FDTD) technique, focusing on boundary and medium modeling. A family of nonstaggered 3-D compact explicit FDTD schemes is analyzed in terms of stability, accuracy, and computational efficiency, and the most accurate and isotropic schemes based on a rectilinear grid are identified. A frequency-dependent boundary model that is consistent with locally reacting surface theory is also presented, in which the wall impedance is represented with a digital filter. For boundaries, accuracy in numerical reflection is analyzed and a stability proof is provided. The results indicate that the proposed 3-D interpolated wideband and isotropic schemes outperform directly related techniques based on Yee's staggered grid and standard digital waveguide mesh, and that the boundary formulations generally have properties that are similar to that of the basic scheme used. Konrad Kowalczyk, Maarten van Walstijn |
IEEE Trans. Speech Audio Process. | 1 |
| 2011 | A Phase Grating Approach to Modeling Surface Diffusion in FDTD Room Acoustics SimulationsabstractIn this paper, a method for modeling diffusive boundaries in finite-difference time-domain (FDTD) room acoustics simulations with the use of impedance filters is presented. The proposed technique is based on the concept of phase grating diffusers, and realized by designing boundary impedance filters from normal-incidence reflection filters with added delay. These added delays, that correspond to the diffuser well depths, are varied across the boundary surface, and implemented using Thiran allpass filters. The proposed method for simulating sound scattering is suitable for modeling high frequency diffusion caused by small variations in surface roughness and, more generally, diffusers characterized by narrow wells with infinitely thin separators. This concept is also applicable to other wave-based modeling techniques. The approach is validated by comparing numerical results for Schroeder diffusers to measured data. In addition, it is proposed that irregular surfaces are modeled by shaping them with Brownian noise, giving good control over the sound scattering properties of the simulated boundary through two parameters, namely the spectral density exponent and the maximum well depth. Konrad Kowalczyk, Maarten van Walstijn, Damian T. Murphy |
IEEE Trans. Speech Audio Process. | 1 |
| 2010 | A comparison of nonstaggered compact FDTD schemes for the 3D wave equationabstractThis paper aims at providing a better insight into the 3D approximations of the wave equation using compact finite-difference time-domain (FDTD) schemes in the context of room acoustic simulations. A general family of 3D compact explicit and implicit schemes based on a nonstaggered rectilinear grid is analyzed in terms of stability, numerical error, and accuracy. Various special cases are compared and the most accurate explicit and implicit schemes are identified. Further considerations presented in the paper include the direct relationship with other numerical approaches found in the literature on room acoustic modeling such as the 3D digital waveguide mesh and Yee's staggered grid technique. Konrad Kowalczyk, Maarten van Walstijn |
ICASSP | 1 |
| 2010 | Wideband and Isotropic Room Acoustics Simulation Using 2-D Interpolated FDTD SchemesabstractIn this paper, a complete method for finite-difference time-domain modeling of rooms in 2-D using compact explicit schemes is presented. A family of interpolated schemes using a rectilinear, nonstaggered grid is reviewed, and the most accurate and isotropic schemes are identified. Frequency-dependent boundaries are modeled using a digital impedance filter formulation that is consistent with locally reacting surface theory. A structurally stable and efficient boundary formulation is constructed by carefully combining the boundary condition with the interpolated scheme. An analytic prediction formula for the effective numerical reflectance is given, and a stability proof provided. The results indicate that the identified accurate and isotropic schemes are also very accurate in terms of numerical boundary reflectance, and outperform directly related methods such as Yee's scheme and the standard digital waveguide mesh. In addition, one particular scheme-referred to here as the interpolated wideband scheme-is suggested as the best scheme for most applications. Konrad Kowalczyk, Maarten van Walstijn |
IEEE Trans. Speech Audio Process. | 1 |
| 2007 | On-Line Simulation of 2D Resonators with Reduced Dispersion Error using Compact Implicit Finite Difference MethodsabstractThis paper presents a method for on-line simulation of 2D resonators with reduced direction-dependent frequency error. The use of a compact implicit finite difference (FD) technique is proposed to reduce the dispersion error remarkably. In particular, a computationally efficient method that allows solving 2D implicit problems with a set of three-diagonal equations, namely the alternating direction implicit is discussed. Efficient equation factorisation together with optimally matched free parameters allows more accurate simulation for wider frequency ranges. With the use of this technique, the dispersion error is limited to 1.1% within the bandwidth up to half of the Nyquist frequency. The compact implicit scheme is compared to compact explicit FD schemes in terms of numerical dispersion error, membrane impulse response, and computational cost. Konrad Kowalczyk, Maarten van Walstijn |
ICASSP (1) | 1 |