VLDB 2026 Research / reviewers in the wild / expert
Mieszko Fras
dblp:296/4461
· DBLP profile ↗
8ranked-venue papers
7as first author
8since 2021 · last 2024
0000-0001-9176-8792ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 6 first-author · 7 since 2021Artificial intelligence and machine learning · 4 · 4 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Reverberant Source Separation Using NTF With Delayed Subsources and Spatial PriorsabstractSpeech signals recorded by distant microphones are often contaminated with room reverberation and signals of interfering speakers. This article addresses the problem of joint source separation and dereverberation using multichannel nonnegative tensor factorization (NTF) in which late reverberant components are modeled using the so-called delayed subsources. The article formulates two distinct signal models of the time-frequency spectrum of the multichannel microphone mixture, in which reverberation is modeled either independently for each source using delayed source variances or jointly using delayed microphone signals. In addition, it defines computationally efficient variants of these two methods with a simplified spatial model in which spatial properties of the late reverberant components are estimated jointly for all delays. For each of the four distinct algorithms, the article first formulates a maximum a posteriori (MaP) estimator based on the NTF model with the localization prior over the mixing matrix that is suitable for the estimation of the early reverberation (primarily the direct-path) signals in a reverberant environment. Next it derives update equations for the four resulting expectation-maximization algorithms, which are thoroughly evaluated and shown to outperform similar state-of-the-art approaches. The results of experimental evaluations, performed using real and simulated data, for determined, over-determined and under-determined scenarios, indicate superior performance of the proposed processing over state-of-the-art in terms of standard source separation and dereverberation metrics. Mieszko Fras, Konrad Kowalczyk |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2023 | Joint Blind Source Separation and Dereverberation for Automatic Speech Recognition using Delayed-Subsource MNMF with Localization Prior
Mieszko Fras, Marcin Witkowski, Konrad Kowalczyk |
INTERSPEECH | 1 |
| 2022 | Convolutional Weighted Minimum Mean Square Error Filter for Joint Source Separation and DereverberationabstractPractical scenarios with multiple simultaneously active speakers recorded using one or more microphones in reverberant rooms pose a challenging problem when the extraction of the desired speaker signal is sought for. The majority of techniques found in the literature facilitate either source separation or dereverberation, which can at best be performed as subsequent, cascade processing. Recently, a solution to the joint task has been proposed, which is known as the weighted power minimization distortionless response (WPD) beamformer. In this paper, we derive a convolutional multichannel filter which performs jointly optimum dereverberation and desired source signal extraction. We formulate a single optimization criterion which minimizes the convolutional source-variance weighted mean square error (CW-MMSE), thereby effectively unifying the weighted prediction error (WPE) based dereverberation and MMSE filtering for the desired source extraction from reverberant mixtures of speakers. Experimental results show a significant performance improvement over the compared state-of-the-art methods such as WPD for datasets with simulated and recorded impulse responses. Mieszko Fras, Marcin Witkowski, Konrad Kowalczyk |
ICASSP | 1 |
| 2022 | Convolutive Weighted Multichannel Wiener Filter Front-end for Distant Automatic Speech Recognition in Reverberant Multispeaker Scenarios
Mieszko Fras, Marcin Witkowski, Konrad Kowalczyk |
INTERSPEECH | 1 |
| 2022 | Convolutional Weighted Parametric Multichannel Wiener Filter for Reverberant Source SeparationabstractIn this letter, we address the problem of simultaneous separation and dereverberation of overlapped speech recorded in reverberant conditions. The majority of state-of-the-art techniques are tailored for either of the two problems, which for the joint task leads to sub-optimum performance or solutions which involve subsequent, cascade processing. In contrast, we propose a jointly optimum approach in which we formulate a single optimization criterion that minimizes variance of undesired signal components at the output of a convolutional filter, weighted with the desired speech variance, subject to a constraint which allows to control the amount of distortions in the estimated speech signal. We then derive a closed-form solution of the proposed convolutional weighted parametric multichannel Wiener (CW-PMW) filter which integrates linear-prediction based dereverberation and speech-distortion weighted Wiener filtering in a jointly optimum manner. The results of experiments performed using measured and simulated data indicate superior performance of the proposed approach in comparison with state-of-the-art, which includes sub-optimum cascades of optimum filters for individual tasks, as well as the recently presented jointly optimum, weighted power minimization distortionless response (WPD) beamformer. Mieszko Fras, Konrad Kowalczyk |
IEEE Signal Process. Lett. | 1 |
| 2021 | Maximum a Posteriori Estimator for Convolutive Sound Source Separation with Sub-Source Based NTF Model and the Localization Probabilistic Prior on the Mixing MatrixabstractIn this paper we present a method for the separation of sound source signals recorded using multiple microphones in a reverberant room. In particular, we propose a maximum a posteriori (MAP) estimator based on the multichannel nonnegative tensor factorization (NTF) model with the localization prior distribution on the mixing matrix, in which the latent data consists of the so-called sub-sources for an improved performance in a reverberant environment. For the proposed MAP estimator, we derive the sub-source based expectation maximization (EM) algorithm with the multiplicative update rules (MU) and the localization prior distribution (LP) on the mixing matrix (SSEM-MU-LP). We then perform several experiments for speech and instrumental sound sources recorded using two microphones, in determined and under-determined scenarios, and with different types of initialization of the model parameters. The results of these experiments clearly indicate a significant improvement of the proposed algorithm with the localization prior over the state-of-the-art NTF-based source separation algorithms, which can reach up to 50% in the signal-to-distortion ratio. Mieszko Fras, Konrad Kowalczyk |
ICASSP | 1 |
| 2021 | Combating Reverberation in NTF-Based Speech Separation Using a Sub-Source Weighted Multichannel Wiener Filter and Linear Prediction
Mieszko Fras, Marcin Witkowski, Konrad Kowalczyk |
Interspeech | 1 |
| 2021 | Incorporation of Localization Information for Sound Source Separation in Spherical Harmonic DomainabstractThis paper concerns the problem of convolutive sound source separation from mutlichannel recordings made with a spherical microphone array. In particular, we formulate two state-of-the-art separation techniques based on Expectation Maximization (EM) and Nonnegative Tensor Factorization (NTF) in the spherical harmonic domain (SHD). Furthermore, we adjust and incorporate the Gaussian Localization Prior (GLP) to the proposed algorithms, which yields two variants of the derived methods. For the source signal reconstruction, a Minimum Variance Distortionless Response (MVDR) beamformer with a single-channel Wiener post-filter is employed. The performance comparison is based on experimental evaluation using micro-phone signals simulated with the image-source method in several scenarios, including diverse geometrical setup, different number of sources and various types of source signals, namely the recordings of speech utterances and musical instruments. The experimental results for the first-order ambisonic signals show that the proposed methods enable high-quality sound source separation in the spherical harmonic domain. In particular, we show that incorporation of the Gaussian Localization Prior to the proposed algorithms leads to a substantial improvement in separation performance. Mateusz Guzik, Mieszko Fras, Konrad Kowalczyk |
MMSP | 2 |