EDBT 2026 Demo / reviewers in the wild / expert
Li Li 0063
dblp:53/2189-63
· DBLP profile ↗
19ranked-venue papers
10as first author
10since 2021 · last 2025
0000-0002-3121-7857ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 9 first-author · 8 since 2021Artificial intelligence and machine learning · 9 · 5 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | First Analyze Then Enhance: A Task-Aware System for Speech Separation, Denoising, and Dereverberation
Shaoxiang Dang, Li Li 0063, Shogo Seki, Hiroaki Kudo |
INTERSPEECH | 2 |
| 2024 | Remixed2remixed: Domain Adaptation for Speech Enhancement by Noise2noise Learning with RemixingabstractThis paper proposes a domain adaptation method for speech enhancement called Remixed2Remixed. The proposed method adopts Noise2Noise (N2N) learning to adapt models trained on artificially generated (out-of-domain: OOD) noisy-clean pairs of data to better separate real-world recorded (in-domain) noisy data. The proposed method employs a teacher model trained on OOD data to acquire pseudo-in-domain speech and noise signals, which are shuffled and remixed twice in each batch to generate two bootstrapped mixtures. The student model is then trained by optimizing an N2N-based cost function computed using these two bootstrapped mixtures. As the training strategy is similar to that of the recently proposed RemixIT, we also investigate the effectiveness of the N2N-based loss as a regularization of RemixIT. Experimental results on the CHiME-7 unsupervised domain adaptation for conversational speech enhancement (UDASE) task revealed that the proposed method outperformed the challenging baseline system, RemixIT, and reduced the performance blurring caused by the teacher models. Li Li 0063, Shogo Seki |
ICASSP | 1 |
| 2024 | Improved Remixing Process for Domain Adaptation-Based Speech Enhancement by Mitigating Data Imbalance in Signal-to-Noise Ratio
Li Li 0063, Shogo Seki |
INTERSPEECH | 1 |
| 2024 | Dual-Channel Target Speaker Extraction Based on Conditional Variational Autoencoder and Directional InformationabstractTarget speaker extraction (TSE) has become an attractive research topic in recent years. However, TSE under the underdetermined conditions is still a challenge. In this paper, we deal with a dual-channel TSE problem under underdetermined conditions. Geometric source separation (GSS) is used to be a solution to the TSE problem, but the performance of conventional GSS methods is limited under underdetermined conditions because of the lack of a powerful source model. We propose a dual-channel TSE method with the combined capabilities of target selection based on geometric constraints, more powerful source modeling, and nonlinear postprocessing. A geometric constraint (GC) on the target direction of arrival (DOA) is applied to select the target, and two conditional variational autoencoders (CVAEs) are used to model a single speaker's speech and interference mixture speech. For postprocessing, an ideal ratio time–frequency (T–F) mask estimated from the separated interference mixture speech is used to extract the target speaker's speech. Moreover, to overcome the impact of DOA estimation errors, we improve the objective function so that the target DOA information can be modified. The experimental results demonstrate that the proposed method achieves 6.24 dB and 8.37 dB improvements compared with the baseline method in terms of signal-to-distortion ratio (SDR) and source-to-interference ratio (SIR), respectively, under medium reverberation for 470 ms. Furthermore, through the analysis of experimental results, we found that the improvement method is robust against DOA estimation errors. Rui Wang 0169, Li Li 0063, Tomoki Toda |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2023 | FastMVAE2: On Improving and Accelerating the Fast Variational Autoencoder-Based Source Separation Algorithm for Determined MixturesabstractThis article proposes a new source model and training scheme to improve the accuracy and speed of the multichannel variational autoencoder (MVAE) method. The MVAE method is a recently proposed powerful multichannel source separation method. It consists of pretraining a source model represented by a conditional VAE (CVAE) and then estimating separation matrices along with other unknown parameters so that the log-likelihood is non-decreasing given an observed mixture signal. Although the MVAE method has been shown to provide high source separation performance, one drawback is the computational cost of the backpropagation steps in the separation-matrix estimation algorithm. To overcome this drawback, a method called “FastMVAE” was subsequently proposed, which uses an auxiliary classifier VAE (ACVAE) to train the source model. By using the classifier and encoder trained in this way, the optimal parameters of the source model can be inferred efficiently, albeit approximately, in each step of the algorithm. However, the generalization capability of the trained ACVAE source model was not satisfactory, which led to poor performance in situations with unseen data. To improve the generalization capability, this article proposes a new model architecture (called the “ChimeraACVAE” model) and a training scheme based on knowledge distillation. The experimental results revealed that the proposed source model trained with the proposed loss function achieved better source separation performance with less computation time than FastMVAE. We also confirmed that our methods were able to separate 18 sources with a reasonably good accuracy. Li Li 0063, Hirokazu Kameoka, Shoji Makino |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2022 | Attentionpit: Soft Permutation Invariant Training for Audio Source Separation with Attention MechanismabstractPermutation invariant training (PIT) has recently attracted attention as a framework to achieve end-to-end time-domain audio source separation. Its goal is to train a separation network that takes a mixture signal as input and produces the J underlying source signals. Since the order of the output signals is arbitrary, the idea of PIT is to first find the best output-target assignment and then update the network parameters based on the error given by that assignment at each iteration. However, there are two known problems with PIT: One is that it has a time complexity of $\mathcal{O}\left( {J!} \right)$, which makes it infeasible as J increases, and the other is that it is prone to getting stuck in bad local optimal solutions due to the hard output-target assignment process. To overcome these problems simultaneously, in this paper, we propose AttentionPIT, which uses an attention mechanism to find soft output-target assignments for separation network training, and can be run in polynomial time in J, as with the recently proposed fast PIT variants such as SinkPIT and HungarianPIT. The training loss of AttentionPIT is fully differentiable, allowing us to simultaneously perform processes corresponding to soft output-target assignment and network parameter update through backpropagation. Experiments on the LibriMix corpus revealed that while AttentionPIT works reasonably well on its own, it works even better when combined with SinkPIT and HungarianPIT so that AttentionPIT is run only in the early stages of training. Hirokazu Kameoka, Shogo Seki, Li Li 0063, Chihiro Watanabe |
ICASSP | 3 |
| 2022 | HBP: An Efficient Block Permutation Solver Using Hungarian Algorithm and Spectrogram Inpainting for Multichannel Audio Source SeparationabstractThis paper proposes a method called "Hungarian Block Permutation (HBP)" to solve the block permutation problem in frequency-domain multichannel audio source separation. Many methods for frequency-domain multichannel audio source separation are designed to simultaneously solve frequency-wise source separation and permutation alignment in determined cases. However, in practice, separation can fail due to permutation inconsistencies in different frequency blocks for various reasons, such as convergence to a locally optimal solution as a result of bad initialization. To correct permutation inconsistencies, the proposed HBP method first masks, for each separated signal, the frequency bands where the components from other sources are likely to be dominant, and then restores the components in those bands so that the restored spectrogram becomes closer to the original spectrogram of the corresponding source. The Hungarian algorithm is then used to perform permutation realignment in those bands in accordance with the restored spectrogram. The experimental results show that the proposed method can solve the permutation realignment and improve the separation performance even in the case of 18 speakers. Li Li 0063, Hirokazu Kameoka, Shogo Seki |
ICASSP | 1 |
| 2022 | Investigation And Comparison of Optimization Methods for Variational Autoencoder-Based Underdetermined Multichannel Source SeparationabstractIn this paper, we investigate two algorithms for variational autoencoder (VAE)-based underdetermined multichannel source separation. We previously extended the multichannel VAE (MVAE) method for determined multichannel source separation and proposed the generalized MVAE (GMVAE) method for underdetermined multichannel source separation. The GMVAE method employs a conditional VAE (CVAE) as the source model representing the power spectrograms of the underlying sources present in a mixture. While we developed a convergence-guaranteed parameter estimation algorithm using a majorization-minimization/minorization-maximization (MM) algorithm, an expectation-maximization (EM) algorithm also allows us to design another algorithm with the same property. However, a comparison of the MM-based and EM-based algorithms has not yet been revealed. To elucidate this, we investigate the MM-based and EM-based algorithms for the GMVAE method, using an improved CVAE variant called auxiliary classifier VAE (ACVAE). The experimental results suggest that the EM-based algorithm takes less computational cost, achieving comparable separation performance with the MM-based algorithm. Shogo Seki, Hirokazu Kameoka, Li Li 0063 |
ICASSP | 3 |
| 2021 | SepNet: A Deep Separation Matrix Prediction Network for Multichannel Audio Source SeparationabstractIn this paper, we propose SepNet, a deep neural network (DNN) designed to predict separation matrices from multichannel observations. One well-known approach to blind source separation (BSS) involves independent component analysis (ICA). A recently developed method called independent low-rank matrix analysis (ILRMA) is one of its powerful variants. These methods allow the estimation of separation matrices based on deterministic iterative algorithms. Specifically, ILRMA is designed to update the separation matrix according to an update rule derived based on the majorization-minimization principle. Although ILRMA performs reasonably well under some conditions, there is still room for improvement in terms of both separation accuracy and computation time, especially for large-scale microphone arrays. The existence of a deterministic iterative algorithm that can find one of the stationary points of the BSS problem implies that a DNN can also play that role if designed and trained properly. Motivated by this, we propose introducing a DNN that learns to convert a predefined input (e.g., an identity matrix) into a true separation matrix in accordance with a multichannel observation. To enable it to find one of the multiple solutions corresponding to different permutations of the source indices, we further propose adopting a permutation invariant training strategy to train the network. By using a fully convolutional architecture, we can design the network so that the forward propagation can be computed efficiently. The experimental results revealed that SepNet was able to find separation matrices faster and with better separation accuracy than ILRMA for mixtures of two sources. Shota Inoue, Hirokazu Kameoka, Li Li 0063, Shoji Makino |
ICASSP | 3 |
| 2021 | Teacher-Student Learning for Low-Latency Online Speech Enhancement Using Wave-U-NetabstractIn this paper, we propose a low-latency online extension of wave-U-net for single-channel speech enhancement, which utilizes teacher-student learning to reduce the system latency while keeping the enhancement performance high. Wave-U-net is a recently proposed end-to-end source separation method, which achieved remarkable performance in singing voice separation and speech enhancement tasks. Since the enhancement is performed in the time domain, wave-U-net can efficiently model phase information and address the domain transformation limitation, where the time-frequency domain is normally adopted. In this paper, we apply wave-U-net to face-to-face applications such as hearing aids and in-car communication systems, where a strictly low-latency of less than 10 ms is required. To this end, we investigate online versions of wave-U-net and propose the use of teacher-student learning to prevent the performance degradation caused by the reduction in input segment length such that the system delay in a CPU is less than 10 ms. The experimental results revealed that the proposed model could perform in real-time with low-latency and high performance, achieving a signal-to-distortion ratio improvement of about 8.73 dB. Sotaro Nakaoka, Li Li 0063, Shota Inoue, Shoji Makino |
ICASSP | 2 |
| 2020 | Geometrically Constrained Independent Vector Analysis for Directional Speech EnhancementabstractThis paper addresses the multichannel directional speech enhancement problem with geometrically constrained independent vector analysis (GCIVA), where we aim to combine the high separation performance from blind source separation and the capability of directional focus from beamforming. The proposed method exploits geometric constraints composed from the spatial information of sources to guide the target speech to the desired output channel. A convergence-guaranteed parameter estimation algorithm is derived from the framework of auxiliary function-based IVA (AuxIVA) to take advantage of fast convergence, low computational cost, and no step-size tuning. We propose a dual-microphone speech enhancement system based on the proposed method and investigate its effectiveness with objective metrics. The experimental evaluations revealed that the proposed system outperformed the conventional beamforming and the standard AuxIVA in a large margin in terms of source-to-distortion and source-to-interference ratios. Li Li 0063, Kazuhito Koishida |
ICASSP | 1 |
| 2020 | Online Directional Speech Enhancement Using Geometrically Constrained Independent Vector Analysis
Li Li 0063, Kazuhito Koishida, Shoji Makino |
INTERSPEECH | 1 |
| 2019 | Joint Separation and Dereverberation of Reverberant Mixtures with Multichannel Variational AutoencoderabstractIn this paper, we deal with a multichannel source separation problem under a highly reverberant condition. The multichannel variational autoencoder (MVAE) is a recently proposed source separation method that employs the decoder distribution of a conditional VAE (CVAE) as the generative model for the complex spectrograms of the underlying source signals. Although MVAE is notable in that it can significantly improve the source separation performance compared with conventional methods, its capability to separate highly reverberant mixtures is still limited since MVAE uses an instantaneous mixture model. To overcome this limitation, in this paper we propose extending MVAE to simultaneously solve source separation and dereverberation problems by formulating the separation system as a frequency-domain convolutive mixture model. A convergence-guaranteed algorithm based on the coordinate descent method is derived for the optimiza- tion. Experimental results revealed that the proposed method outperformed the conventional methods in terms of all the source separation criteria in highly reverberant environments. Shota Inoue, Hirokazu Kameoka, Li Li 0063, Shogo Seki, Shoji Makino |
ICASSP | 3 |
| 2019 | Fast MVAE: Joint Separation and Classification of Mixed Sources Based on Multichannel Variational Autoencoder with Auxiliary ClassifierabstractThis paper proposes an alternative algorithm for the multi-channel variational autoencoder (MVAE), a recently proposed multichannel source separation approach. While MVAE is notable for its impressive source separation performance, its convergence-guaranteed optimization algorithm and the fact that it allows us to estimate source-class labels simultaneously with source separation, there are still two major drawbacks, namely, the high computational complexity and the unsatisfactory source classification accuracy. To overcome these drawbacks, the proposed method employs an auxiliary classifier VAE, which is an information-theoretic extension of the conditional VAE, for learning the generative model of the source spectrograms. Furthermore, with the trained auxiliary classifier, we introduce a novel algorithm for the optimization that can both reduce the computational time and improve the source classification performance. We call the proposed method "fast MVAE (fMVAE) ". Experimental evaluations revealed that fMVAE achieved source separation performance comparable to that of MVAE and a source classification accu-racy rate of about 80% while reducing computational time by about 93%. Li Li 0063, Hirokazu Kameoka, Shoji Makino |
ICASSP | 1 |
| 2019 | Supervised Determined Source Separation with Multichannel Variational AutoencoderabstractThis letter proposes a multichannel source separation technique, the multichannel variational autoencoder (MVAE) method, which uses a conditional VAE (CVAE) to model and estimate the power spectrograms of the sources in a mixture. By training the CVAE using the spectrograms of training examples with source-class labels, we can use the trained decoder distribution as a universal generative model capable of generating spectrograms conditioned on a specified class index. By treating the latent space variables and the class index as the unknown parameters of this generative model, we can develop a convergence-guaranteed algorithm for supervised determined source separation that consists of iteratively estimating the power spectrograms of the underlying sources, as well as the separation matrices. In experimental evaluations, our MVAE produced better separation performance than a baseline method. Hirokazu Kameoka, Li Li 0063, Shota Inoue, Shoji Makino |
Neural Comput. | 2 |
| 2018 | Deep Clustering with Gated Convolutional NetworksabstractDeep clustering is a recently introduced deep learning-based method for speech separation. The idea is to model and train the mapping from each time-frequency (TF) region of a spectrogram to an embedding space so that the embedding features of the TF regions dominated by the same source are forced to get close to each other and those dominated by different sources are forced to get separated from each other. This allows us to construct binary masks by applying a regular clustering algorithm to the mapped embedding vectors of a test mixture signal. The original deep clustering uses a bidirectional long short-term memory (BLSTM) recurrent neural network (RNN) to model the embedding process. Although RNN-based architectures are indeed a natural choice for modeling long-term dependencies of time series data, recent work has shown that convolutional networks (CNNs) with gating mechanisms also have an excellent potential for capturing long-term structures. In addition, they are less prone to overfitting and are suitable for parallel computations. Motivated by these facts, this paper proposes adopting CNN-based architectures for deep clustering. Specifically, we use a gated CNN architecture, which was introduced to model word sequences for language modeling and was shown to outperform LSTM language models trained in a similar setting. We tested various CNN architectures on a monaural source separation task. The results revealed that the proposed architectures achieved better performance than the BLSTM-based architecture under the same training condition and comparable performance even with a smaller amount of training data. Li Li 0063, Hirokazu Kameoka |
ICASSP | 1 |
| 2018 | Nonnegative Matrix Factorization With Basis Clustering Using Cepstral Distance RegularizationabstractOne successful approach for audio source separation involves applying nonnegative matrix factorization (NMF) to a magnitude spectrogram regarded as a nonnegative matrix. This can be interpreted as approximating the observed spectra at each time frame as the linear sum of the basis spectra scaled by time-varying amplitudes. This paper deals with the problem of the unsupervised instrument-wise source separation of polyphonic signals based on an extension of the NMF approach. We focus on the fact that each piece of music is typically played on a handful of musical instruments, which allows us to assume that the spectra of the underlying audio events in a polyphonic signal can be grouped into a reasonably small number of clusters in the mel-frequency cepstral coefficient (MFCC) domain. Based on this assumption, we propose formulating factorization of a magnitude spectrogram and clustering of the basis spectra in the MFCC domain as a joint optimization problem and derive a novel optimization algorithm based on the majorization–minimization principle. Experimental results revealed that our method was superior to a two-stage algorithm that consists of performing factorization followed by clustering the basis spectra, thus showing the advantage of the joint optimization approach. Hirokazu Kameoka, Takuya Higuchi, Mikihiro Tanaka, Li Li 0063 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2017 | Speech Enhancement Using Non-Negative Spectrogram Models with Mel-Generalized Cepstral Regularization
Li Li 0063, Hirokazu Kameoka, Tomoki Toda, Shoji Makino |
INTERSPEECH | 1 |
| 2016 | Semi-Supervised Joint Enhancement of Spectral and Cepstral Sequences of Noisy Speech
Li Li 0063, Hirokazu Kameoka, Takuya Higuchi, Hiroshi Saruwatari |
INTERSPEECH | 1 |