VLDB 2026 Research / reviewers in the wild / expert
Zbynek Koldovský
dblp:70/1809
· DBLP profile ↗
35ranked-venue papers
16as first author
7since 2021 · last 2025
0000-0002-1791-5675ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 28 · 12 first-author · 6 since 2021Artificial intelligence and machine learning · 12 · 7 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Blind Capon Beamformer Based on Independent Component Extraction: Single-Parameter AlgorithmabstractWe consider a phase-shift mixing model for linear sensor arrays in the context of blind source extraction. We derive a blind Capon beamformer that seeks the direction where the output is independent of the other signals in the mixture. The algorithm is based on Independent Component Extraction and imposes an orthogonal constraint, thanks to which it optimizes only one real-valued parameter related to the angle of arrival. The Cramér-Rao lower bound for the mean interference-to-signal ratio is derived. The algorithm and the bound are compared with conventional blind and direction-of-arrival estimation+beamforming methods, showing improvements in terms of extraction accuracy. An application is demonstrated in frequency-domain speaker extraction in a low-reverberation room. Zbynek Koldovský, Jaroslav Cmejla, Stephen O'Regan |
IEEE Signal Process. Lett. | 1 |
| 2024 | Comparison Of Frequency-Fusion Mechanisms For Binaural Direction-Of-Arrival Estimation For Multiple SpeakersabstractTo estimate the direction of arrival (DOA) of multiple speakers with methods that use prototype transfer functions, frequency-dependent spatial spectra (SPS) are usually constructed. To make the DOA estimation robust, SPS from different frequencies can be combined. According to how the SPS are combined, frequency fusion mechanisms are categorized into narrowband, broadband, or speaker-grouped, where the latter mechanism requires a speaker-wise grouping of frequencies. For a binaural hearing aid setup, in this paper we propose an interaural time difference (ITD)-based speaker-grouped frequency fusion mechanism. By exploiting the DOA dependence of ITDs, frequencies can be grouped according to a common ITD and be used for DOA estimation of the respective speaker. We apply the proposed ITD-based speaker-grouped frequency fusion mechanism for different DOA estimation methods, namely the multiple signal classification, steered response power and a recently published method based on relative transfer function (RTF) vectors. In our experiments, we compare DOA estimation with different fusion mechanisms. For all considered DOA estimation methods, the proposed ITD-based speaker-grouped frequency fusion mechanism results in a higher DOA estimation accuracy compared with the narrowband and broadband fusion mechanisms. Daniel Fejgin, Elior Hadad, Sharon Gannot, Zbynek Koldovský, Simon Doclo |
ICASSP | 4 |
| 2023 | Dynamic Independent Component Extraction with Blending Mixing Vector: Lower Bound on Mean Interference-to-Signal RatioabstractThis paper deals with dynamic Blind Source Extraction (BSE) from where the mixing parameters characterizing the position of a source of interest (SOI) are allowed to vary over time. We present a new source extraction model called CvxCSV which is a parameter-reduced modification of the recent Constant Separation Vector (CSV) mixing model. In CvxCSV, the mixing vector evolves as a convex combination of its initial and final values. We derive a lower bound on the achievable mean interference-to-signal ratio (ISR) based on the Cramér-Rao theory. The bound reveals advantageous properties of CvxCSV compared with CSV and compared with a sequential BSE based on independent component extraction (ICE). In particular, the achievable ISR by CvxCSV is lower than that by the previous approaches. Moreover, the model requires significantly weaker conditions for identifiability, even when the SOI is Gaussian. Jaroslav Cmejla, Zbynek Koldovský, Vaclav Kautsky, Tülay Adali |
ICASSP | 2 |
| 2022 | Target Speech Extraction: Independent Vector Extraction Guided by Supervised Speaker IdentificationabstractThis manuscript proposes a novel robust procedure for the extraction of a speaker of interest (SOI) from a mixture of audio sources. The estimation of the SOI is performed via independent vector extraction (IVE). Since the blind IVE cannot distinguish the target source by itself, it is guided towards the SOI via frame-wise speaker identification based on deep learning. Still, an incorrect speaker can be extracted due to guidance failings, especially when processing challenging data. To identify such cases, we propose a criterion for non-intrusively assessing the estimated speaker. It utilizes the same model as the speaker identification, so no additional training is required. When incorrect extraction is detected, we propose a “deflation” step in which the incorrect source is subtracted from the mixture and, subsequently, another attempt to extract the SOI is performed. The process is repeated until successful extraction is achieved. The proposed procedure is experimentally tested on artificial and real-world datasets containing challenging phenomena: source movements, reverberation, transient noise, or microphone failures. The method is compared with state-of-the-art blind algorithms as well as with current fully supervised deep learning-based methods. Jirí Málek, Jakub Janský, Zbynek Koldovský, Tomás Kounovský, Jaroslav Cmejla, Jindrich Zdánský |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2021 | Blind Extraction of Moving Sources via Independent Component and Vector Analysis: ExamplesabstractThis paper is devoted to the recently proposed mixing model with constant separating vector (CSV) for Blind Source Extraction of moving sources using the FastDIVA algorithm, which is an extension of the famous FastICA and FastIVA for static mixtures. The benefits due to the CSV model and FastDIVA are demonstrated in three new applications. First, the extraction of a moving speaker in a noisy reverberant environment using a dense array of 48 MEMS microphones is considered. Second, a case study on the blind extraction of moving brain activity from visually evoked potentials in electroencephalogram is reported. Third, a simulation of block-by-block online extraction of a moving source is demonstrated. In these examples, the CSV and FastDIVA show their new potential and good performance in handling the blind moving source extraction problem. N. Amor, Jaroslav Cmejla, Vaclav Kautsky, Zbynek Koldovský, Tomás Kounovský |
ICASSP | 4 |
| 2021 | Blind Extraction of Moving Audio Source in a Challenging Environment Supported by Speaker Identification Via X-VectorsabstractWe propose a novel approach for semi-supervised extraction of a moving audio source of interest (SOI) applicable in reverberant and noisy environments. The blind part of the method is based on independent vector extraction (IVE) and uses the recently proposed constant separating vector (CSV) mixing model. This model allows for changes of mixing parameters within the processed interval of the mixture, which potentially leads to higher accuracy of SOI estimation. The supervised part of the method concerns a pilot signal, which is related to the SOI and ensures the convergence of the blind method towards the SOI. The pilot is based on robust detection of frames where SOI is dominant via speaker embeddings called X-vectors. Robustness of the detection is achieved through augmentation of the data for the supervised training of the X-vectors. The pilot-supported extraction yields significantly better performance compared to its unsupervised counterpart identifying SOI solely using the initialization. Jirí Málek, Jakub Janský, Tomás Kounovský, Zbynek Koldovský, Jindrich Zdánský |
ICASSP | 4 |
| 2021 | Advanced Semi-Blind Speaker Extraction and Tracking Implemented in Experimental Device with Revolving Dense Microphone Array
Jaroslav Cmejla, Tomás Kounovský, Jakub Janský, Jirí Málek, M. Rozkovec, Zbynek Koldovský |
Interspeech | 6 |
| 2020 | Adaptive Blind Audio Source Extraction Supervised By Dominant Speaker Identification Using X-VectorsabstractWe propose a novel algorithm for adaptive blind audio source extraction. The proposed method is based on independent vector analysis and utilizes the auxiliary function optimization to achieve high convergence speed. The algorithm is partially supervised by a pilot signal related to the source of interest (SOI), which ensures that the method correctly extracts the utterance of the desired speaker. The pilot is based on the identification of a dominant speaker in the mixture using x-vectors. The properties of the x-vectors computed in the presence of cross-talk are experimentally analyzed. The proposed approach is verified in a scenario with a moving SOI, static interfering speaker and environmental noise. Jakub Janský, Jirí Málek, Jaroslav Cmejla, Tomás Kounovský, Zbynek Koldovský, Jindrich Zdánský |
ICASSP | 5 |
| 2020 | Block-online multi-channel speech enhancement using deep neural network-supported relative transfer function estimatesabstractThis work addresses the problem of block‐online processing for multi‐channel speech enhancement. Such processing is vital in scenarios with moving speakers and/or when short utterances are processed, e.g. in voice assistant applications. We consider several variants of a system that performs beamforming supported by deep neural network‐based voice activity detection followed by post‐filtering. The speaker is targeted through estimating relative transfer functions between microphones. Each block of the input signals is processed independently to make the method applicable in highly dynamic environments. Due to short processed blocks, the statistics required by the beamformer are estimated less precisely. The influence of this inaccuracy is studied and compared to batch processing regime, when recordings are treated as one block. The experimental evaluation is performed on large datasets of CHiME‐4 and another dataset featuring moving target speaker. The experiments are evaluated in terms of objective and perceptual criteria. Moreover, word error rate (WER) of a speech recognition system is evaluated, for which the method serves as a front‐end. The results indicate that the proposed method is robust for short length of the processed block. Significant improvements in terms of the criteria and WER are observed even for the block length of 250 ms. Jirí Málek, Zbynek Koldovský, Marek Bohac |
IET Signal Process. | 2 |
| 2019 | Performance Bound for Blind Extraction of Non-gaussian Complex-valued Vector Component from Gaussian BackgroundabstractIndependent Vector Extraction aims at the joint blind source extraction of K dependent signals of interest (SOI) from K mixtures (one signal from one mixture). Similarly to Independent Component/Vector Analysis (ICA/IVA), the SOIs are assumed to be independent of the other signals in the mixture. Compared to IVA, the (de-)mixing IVE model is reduced in the number of parameters for the extraction problem. The SOIs are assumed to be non-Gaussian or noncircular Gaussian, while the other signals are modeled as circular Gaussian. In this paper, a Cramér-Rao-Induced Bound (CRIB) for the achievable Interference-to-Signal Ratio (ISR) is derived for IVE. The bound is compared with similar bounds for ICA, IVA, and Independent Component Extraction (ICE). Numerical simulations show a good correspondence between the empirical results and the theory. Vaclav Kautsky, Zbynek Koldovský, Petr Tichavský |
ICASSP | 2 |
| 2019 | Extraction of Independent Vector Component from Underdetermined Mixtures through Block-wise Determined ModelingabstractWe propose a new model for blind source extraction where the source of interest is assumed to be static while the background noise is dynamic. The model is determined within short blocks (the same number of sources as that of sensors), however, the noise subspace can be changing from block to block. We propose a gradient-based algorithm that jointly extracts an independent vector component from a set of mixtures obeying the model based on maximum quasi-likelihood principle. Simulations confirm the validity of the approach, and experiments with real-world recordings show promising results. Zbynek Koldovský, Jirí Málek, Jakub Janský |
ICASSP | 1 |
| 2017 | Supervised independent vector analysis through pilot dependent componentsabstractUnknown global permutation of the separated sources, time-varying source activity and under determination are common problems affecting on-line Independent Vector Analysis when applied to real-world speech enhancement. In this work we propose to extend the signal model of IVA by introducing additional supervising components. Pilot signals, which are dependent on the sources, are injected in the multidimensional source representation and act as a prior knowledge. The resulting adaptation still maximizes the multivariate source independence, while simultaneously forcing the estimation of sources dependent on the pilot components. It is also shown as the S-IVA is a generalization over the previously proposed weighted Natural Gradient. Numerical evaluations shows the effectiveness of the proposed method in challenging real-world applications. Francesco Nesta, Zbynek Koldovský |
ICASSP | 2 |
| 2016 | Blind separation of underdetermined linear mixtures based on source nonstationarity and AR(1) modelingabstractThe problem of blind separation of underdetermined instantaneous mixtures of independent signals is addressed through a method relying on nonstationarity of the original signals. The signals are assumed to be piecewise stationary with varying variances in different epochs. In comparison with previous works, in this paper it is assumed that the signals are not i.i.d. in each epoch, but obey a first-order autoregressive model. This model was shown to be more appropriate for blind separation of natural speech signals. A separation method is proposed that is nearly statistically efficient (approaching the corresponding Cramér-Rao lower bound), if the separated signals obey the assumed model. In the case of natural speech signals, the method is shown to have separation accuracy better than the state-of-the-art methods. Ondrej Sembera, Petr Tichavský, Zbynek Koldovský |
ICASSP | 3 |
| 2015 | Spatial Source Subtraction Based on Incomplete Measurements of Relative Transfer FunctionabstractRelative impulse responses between microphones are usually long and dense due to the reverberant acoustic environment. Estimating them from short and noisy recordings poses a long-standing challenge of audio signal processing. In this paper, we apply a novel strategy based on ideas of compressed sensing. Relative transfer function (RTF) corresponding to the relative impulse response can often be estimated accurately from noisy data but only for certain frequencies. This means that often only an incomplete measurement of the RTF is available. A complete RTF estimate can be obtained through finding its sparsest representation in the time-domain: that is, through computing the sparsest among the corresponding relative impulse responses. Based on this approach, we propose to estimate the RTF from noisy data in three steps. First, the RTF is estimated using any conventional method such as the nonstationarity-based estimator by Gannotor through blind source separation. Second, frequencies are determined for which the RTF estimate appears to be accurate. Third, the RTF is reconstructed through solving a weighted${\ell _1}$convex program, which we propose to solve via a computationally efficient variant of the SpaRSA (Sparse Reconstruction by Separable Approximation) algorithm. An extensive experimental study with real-world recordings has been conducted. It has been shown that the proposed method is capable of improving many conventional estimators used as the first step in most situations. Zbynek Koldovský, Jirí Málek, Sharon Gannot |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2014 | A homotopy recursive-in-model-order algorithm for weighted LassoabstractA fast algorithm to solve weighted ℓ1-minimization problems with N × N square “measuring” matrices is proposed. The method is recursive-in-model-order and tracks a homotopy path that goes through solutions of the optimization sub-tasks in the order of 1 through N. It thus yields solutions for all model orders and performs this task faster than the other compared methods. We show applications of this method in sparse linear system identification, in particular, the estimation of sparse target-cancellation filters for audio source separation. Zbynek Koldovský, Petr Tichavský |
ICASSP | 1 |
| 2014 | Sparse target cancellation filters with application to semi-blind noise extractionabstractImpulse responses of filters that perform spatial null in a target direction, so-called target-cancellation filters (CFs), are usually long and dense due to the reverberant acoustic environment. It is therefore hard to blindly estimate them from noisy recordings of the target. In this paper, we show that efficient sparse CFs having many coefficients equal to zero can be designed such that their cancellation performance is tolerably lower than the performance of dense CFs. We show that an efficient sparse CF can be blindly estimated from noisy data, provided that its support is known. The resulting filter is better than a dense CF which has been blindly estimated without any prior knowledge. Jirí Málek, Zbynek Koldovský |
ICASSP | 2 |
| 2013 | Noise reduction in dual-microphone mobile phones using a bank of pre-measured target-cancellation filtersabstractIn this paper, a novel method of noise reduction for dual-microphone mobile phones is proposed. The method is based on a set (bank) of target-cancellation filters derived in a noise-free situation for different possible positions of the phone with respect to the speaker mouth. Next, a novel construction of the target-cancellation filter is proposed, which is suitable for the application. The set of the cancellation filters is used to accurately estimate the noise of the environment, which is then subtracted from the recorded signal via standard Wiener filter or a power level difference method. Experiments with recorded data show a good performance and low complexity of the system, making it possible for an integration into mobile communication devices. Zbynek Koldovský, Petr Tichavský, David Botka |
ICASSP | 1 |
| 2013 | A Two-Stage MMSE Beamformer for Underdetermined Signal SeparationabstractBlind separation of underdetermined instantaneous mixtures is a popular solution to inverse problems encountered in audio or biomedical applications where the number of sources exceeds the number of sensors. There are two non-equivalent tasks: to identify the mixing matrix and to separate the original sources. In this paper, we focus on the latter task by proposing a novel beamformer that minimizes the theoretical mean square error distance between the separated and original signals. The beamformer has two stages: one for the estimation of signals and one for their refinement. Within the former stage, the signals are assumed to be random and locally stationary, while the latter stage is based on a semi-deterministic model. The experiments prove superior performance of the proposed method compared to conventional MMSE beamforming. Zbynek Koldovský, Petr Tichavský, Anh Huy Phan 0001, Andrzej Cichocki |
IEEE Signal Process. Lett. | 1 |
| 2013 | Semi-Blind Noise Extraction Using Partially Known Position of the Target SourceabstractAn extracted noise signal provides important information for subsequent enhancement of a target signal. When the target's position is fixed, the noise extractor could be a target-cancellation filter derived in a noise-free situation. In this paper we consider a situation when such cancellation filters are prepared for a set of several possible positions of the target in advance. The set of filters is interpreted as prior information available for the noise extraction when the target's exact position is unknown. Our novel method looks for a linear combination of the prepared filters via Independent Component Analysis. The method yields a filter that has a better cancellation performance than the individual filters or filters based on a minimum variance principle. The method is tested in a highly noisy and reverberant real-world environment with moving target source and interferers. A post-processing by Wiener filter using the noise signal extracted by the method is able to improve signal-to-noise ratio of the target by up to 8 dB. Zbynek Koldovský, Jirí Málek, Petr Tichavský, Francesco Nesta |
IEEE Trans. Speech Audio Process. | 1 |
| 2012 | Low-rank blind nonnegative matrix deconvolutionabstractA novel blind deconvolution is proposed to seek for basis patterns and their location maps inside a nonnegative data matrix. Basis patterns can have different sizes, and shift in independent directions. Moreover, the location maps can be low-rank or rank-one matrices composed by two relatively small and tall matrices or by two vectors. A general framework to solve this problem together with algorithms are introduced. The experiments on music and texture decomposition will confirm performance of our method, and of the proposed algorithms. Anh Huy Phan 0001, Petr Tichavský, Andrzej Cichocki, Zbynek Koldovský |
ICASSP | 4 |
| 2011 | Stability of CANDECOMP-PARAFAC tensor decompositionabstractIn this paper, stability of the CANDECOMP-PARAFAC (CP) tensor decomposition is addressed. It is done by deriving the Cramér-Rao lower bound (CRLB) on variance of an unbiased estimate of the tensor parameters, i.e. elements of its factor matrices, from its noisy observation (the tensor plus a random Gaussian i.i.d. tensor). The existence of the bound reveals necessary conditions for essential uniqueness of the CP decomposition, moreover, for identifiability of each column of each factor matrix separately. Analytical closed-form expressions of the bound are derived for 3 way tensors of rank 1 and 2. As a byproduct, a novel computationally efficient expression for the inverse of the approximate Hessian matrix is derived. Petr Tichavský, Zbynek Koldovský |
ICASSP | 2 |
| 2011 | Blind Speech Separation in Time-Domain Using Block-Toeplitz Structure of Reconstructed Signal MatricesabstractMethods for Blind Source Separation (BSS) aim at recovering signals from their mixture without prior knowledge about the signals and the mixing system. Among others, they provide tools for enhancing speech signals when they are disturbed by unknown noise or other interfering signals in the mixture. This paper considers a recent time-domain BSS method that is based on a complete decomposition of a signal subspace into components that should be independent. The components are used to reconstruct images of original signals using an ad hoc weighting, which influences the final performance of the method markedly. We propose a novel weighting scheme that utilizes block-Toeplitz structure of signal matrices and relies thus on an established property. We provide experiments with blind speech separation and speech recognition that prove the better performance of the modified BSS method. Zbynek Koldovský, Jirí Málek, Petr Tichavský |
INTERSPEECH | 1 |
| 2011 | Time-Domain Blind Separation of Audio Sources on the Basis of a Complete ICA Decomposition of an Observation SpaceabstractTime-domain algorithms for blind separation of audio sources can be classified as being based either on a partial or complete decomposition of an observation space. The decomposition, especially the complete one, is mostly done under a constraint to reduce the computational burden. However, this constraint potentially restricts the performance. The authors propose a novel time-domain algorithm that is based on a complete unconstrained decomposition of the observation space. The observation space may be defined in a general way, which allows application of long separating filters, although its dimension is low. The decomposition is done by an appropriate independent component analysis (ICA) algorithm giving independent components that are grouped into clusters corresponding to the original sources. Components of the clusters are combined by a reconstruction procedure after estimating microphone responses of the original sources. The authors demonstrate by experiments that the method works effectively with short data, compared to other methods. Zbynek Koldovský, Petr Tichavský |
IEEE Trans. Speech Audio Process. | 1 |
| 2010 | Simultaneous search for all modes in multilinear modelsabstractParallel factor (PARAFAC) analysis is an extension of a low rank decomposition to higher way arrays, usually called tensors. Most of existing methods are based on an alternating least square (ALS) algorithm that proceeds iteratively, and minimizes a criterion (that is usually quadratic) of the fit with respect to individual factors one by one. Convergence of this approach is known to be slow, if some of the factor contain nearly co-linear vectors. This problem can be partly alleviated by an enhanced line search (ELS) by Rajih et al. (2008). In this paper we show that the method originally proposed by Paatero (1997), consisting in optimization with respect to all modes simultaneously, can be simplified, and can far outperform the ALS-ELS in ill-conditioned data in all modes. Petr Tichavský, Zbynek Koldovský |
ICASSP | 2 |
| 2009 | A fast asymptotically efficient algorithm for blind separation of a linear mixture of block-wise stationary autoregressive processesabstractWe propose a novel blind source separation algorithm called Block AutoRegressive Blind Identification (BARBI). The algorithm is asymptotically efficient in separation of instantaneous linear mixtures of blockwise stationary Gaussian autoregressive processes. A novel closed-form formula is derived for a Cramér Rao lower bound on elements of the corresponding Interference-to-Signal Ratio (ISR) matrix. This theoretical ISR matrix can serve as an estimate of the separation performance on the particular data. In simulations, the algorithm is shown to be applicable in blind separation of a linear mixture of speech signals. Petr Tichavský, Arie Yeredor, Zbynek Koldovský |
ICASSP | 3 |
| 2009 | Blind separation of piecewise stationary non-Gaussian sources
Zbynek Koldovský, Jirí Málek, Petr Tichavský, Yannick Deville, Shahram Hosseini |
Signal Process. | 1 |
| 2008 | Extension of EFICA algorithm for blind separation of piecewise stationary non Gaussian sourcesabstractWe propose an extension of EFICA algorithm for piecewise stationary and non Gaussian signals. The proposed method is able to profit from varying distribution of the original signals and also from their varying variance, which is demonstrated by simulations with real-world signals. We show that in case of constant-variance signals, the accuracy of the method may achieve the corresponding Cramer-Rao bound, if score functions of the original signals are known in all blocks. Zbynek Koldovský, Jirí Málek, Petr Tichavský, Yannick Deville, Shahram Hosseini |
ICASSP | 1 |
| 2008 | Enhancement of noisy speech recordings via blind source separationabstractWe propose an improved time-domain Blind Source Separa-tion method and apply it to speech signal enhancement using multiple microphone recordings. The improvement consists in utilization of fuzzy clustering instead of a hard one, which is verified by experiments where real-world mixtures of two au-dio signals are separated from two microphones. Performance of the method is demonstrated by recognizing mixed and sepa-rated utterances from the Czech part of the European broadcast news database using our Czech LVCSR system. The separation allows significantly better recognition, e.g., by 32 % when the jammer signal is a Gaussian noise and the input signal-to-noise ratio is 10dB. Jirí Málek, Zbynek Koldovský, Jindrich Zdánský, Jan Nouza |
INTERSPEECH | 2 |
| 2008 | A Hybrid Technique for Blind Separation of Non-Gaussian and Time-Correlated Sources Using a Multicomponent ApproachabstractBlind inversion of a linear and instantaneous mixture of source signals is a problem often encountered in many signal processing applications. Efficient fastICA (EFICA) offers an asymptotically optimal solution to this problem when all of the sources obey a generalized Gaussian distribution, at most one of them is Gaussian, and each is independent and identically distributed (i.i.d.) in time. Likewise, weights-adjusted second-order blind identification (WASOBI) is asymptotically optimal when all the sources are Gaussian and can be modeled as autoregressive (AR) processes with distinct spectra. Nevertheless, real-life mixtures are likely to contain both Gaussian AR and non-Gaussian i.i.d. sources, rendering WASOBI and EFICA severely suboptimal. In this paper, we propose a novel scheme for combining the strengths of EFICA and WASOBI in order to deal with such hybrid mixtures. Simulations show that our approach outperforms competing algorithms designed for separating similar mixtures. Petr Tichavský, Zbynek Koldovský, Arie Yeredor, Germán Gómez-Herrero, Eran Doron |
IEEE Trans. Neural Networks | 2 |
| 2007 | Time-domain blind audio source separation using advanced ICA methodsabstractIn this paper, a prototype of novel algorithm for blind separation of convolutive mixtures of audio sources is proposed. The method works in time-domain, and it is based on the recently very successful algorithm EFICA for Independent Component Analysis, which is an enhanced version of more famous FastICA. Performance of the new algorithm is very promising, at least, comparable to other (mostly frequency domain) algorithms. Audio separation examples are included. Zbynek Koldovský, Petr Tichavský |
INTERSPEECH | 1 |
| 2006 | Methods of Fair Comparison of Performance of Linear ICA Techniques in Presence of Additive NoiseabstractLinear ICA model with additive Gaussian noise is frequently considered in many practical applications, because it approaches the reality often much better than the noise-free alternate. In this paper, a number important differences between noisy and noiseless ICA are discussed. It is shown that estimation of the mixing/demixing matrix should not be the main goal, in the noisy case. Instead, it is proposed to compare outcome of ICA algorithms with a minimum mean square (MMSE) separation, derived for known mixing model. The signal-to-interference-plus-noise ratio is suggested as the most meaningful performance criterion. A simulation study that compares a few well known ICA algorithms applied to noise data is included. Zbynek Koldovský, Petr Tichavský |
ICASSP (5) | 1 |
| 2006 | Continuous time-frequency masking method for blind speech separation with adaptive choice of threshold parameter using ICAabstractWe propose a novel method for blind speech separation using continuous time-frequency masking. The method is equipped with an adaptive choice of a threshold parameter that is based on utilization of ICA methods. We present a direct application that consists in the speech segregation for automatic transcription of spoken broadcasts disturbed by background music. Experimental results show improved performance in comparison with traditionally used binary masking methods. Zbynek Koldovský, Jan Nouza, Jan Kolorenc |
INTERSPEECH | 1 |
| 2006 | Efficient Variant of Algorithm FastICA for Independent Component Analysis Attaining the CramÉr-Rao Lower BoundabstractFastICA is one of the most popular algorithms for independent component analysis (ICA), demixing a set of statistically independent sources that have been mixed linearly. A key question is how accurate the method is for finite data samples. We propose an improved version of the FastICA algorithm which is asymptotically efficient, i.e., its accuracy given by the residual error variance attains the Cramér-Rao lower bound (CRB). The error is thus as small as possible. This result is rigorously proven under the assumption that the probability distribution of the independent signal components belongs to the class of generalized Gaussian (GG) distributions with parameter alpha, denoted GG(alpha) for alpha > 2. We name the algorithm efficient FastICA (EFICA). Computational complexity of a Matlab implementation of the algorithm is shown to be only slightly (about three times) higher than that of the standard symmetric FastICA. Simulations corroborate these claims and show superior performance of the algorithm compared with algorithm JADE of Cardoso and Souloumiac and nonparametric ICA of Boscolo et al. on separating sources with distribution GG (alpha) with arbitrary alpha, as well as on sources with bimodal distribution, and a good performance in separating linearly mixed speech signals. Zbynek Koldovský, Petr Tichavský, Erkki Oja |
IEEE Trans. Neural Networks | 1 |
| 2005 | Cramer-Rao lower bound for linear independent component analysisabstractThis paper derives a closed-form expression for the Cramer-Rao bound (CRB) on estimating the source signals in the linear independent component analysis problem, assuming that all independent components have finite variance. It is also shown that the fixed-point algorithm known as FastICA can approach the CRB (the estimate can be nearly efficient) in two situations: (1) when the distribution of the sources is not too much different from Gaussian, for the symmetric version of the algorithm using any of the custom nonlinear functions (pow3, tanh, gauss); (2) when the distribution of the sources is very different from Gaussian (e.g. has long tails) and the nonlinear function in the algorithm equals the score function of each independent component. Zbynek Koldovský, Petr Tichavský, Erkki Oja |
ICASSP (3) | 1 |
| 2004 | Optimal pairing of signal components separated by blind techniquesabstractIn this letter, the problem of optimal pairing of signal components separated by blind techniques in different time-windows or in different frequency bins is addressed. The optimum pairing is defined as the one which minimizes the sum of some distances (criteria of dissimilarity) of the to-be-assigned signal components. It is shown that the optimal pairing can be achieved by the Kuhn-Munkres algorithm known in graph theory as a solution to the optimal assignment problem. An advantage of the proposed pairing method is shown on data from electroencephalogram, which are blindly separated using the FastICA algorithm in a sliding time-window with the aim to study the time evolution of elements of the estimated mixing matrix. Petr Tichavský, Zbynek Koldovský |
IEEE Signal Process. Lett. | 2 |