EDBT 2026 Demo / reviewers in the wild / expert
Ying Shi 0001
dblp:12/4219-1
· DBLP profile ↗
12ranked-venue papers
5as first author
8since 2021 · last 2026
0009-0005-0834-9615ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 12 · 5 first-author · 8 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Dual-Path Conditional Chain for CTC-Based Multi-Talker Speech Recognition
Ying Shi 0001, Jiqing Han 0001 |
IEEE Signal Process. Lett. | 1 |
| 2025 | Sagalee: an Open Source Automatic Speech Recognition Dataset for Oromo LanguageabstractWe present a novel Automatic Speech Recognition (ASR) dataset for the Oromo language, a widely spoken language in Ethiopia and neighboring regions. The dataset was collected through a crowdsourcing initiative, encompassing a diverse range of speakers and phonetic variations. It consists of 100 hours of real-world audio recordings paired with transcriptions, covering read speech in both clean and noisy environments. This dataset addresses the critical need for ASR resources for the Oromo language which is underrepresented. To show its applicability for the ASR task, we conducted experiments using the Conformer model, achieving a Word Error Rate (WER) of 15.32% with hybrid CTC and AED loss and WER of 18.74% with pure CTC loss. Additionally, fine-tuning the Whisper model resulted in a significantly improved WER of 10.82%. These results establish baselines for Oromo ASR, highlighting both the challenges and the potential for improving ASR performance in Oromo. The dataset is publicly available at https://github.com/turinaf/sagalee and we encourage its use for further research and development in Oromo speech processing. Turi Abu, Ying Shi 0001, Thomas Fang Zheng, Dong Wang 0013 |
ICASSP | 2 |
| 2025 | Knowledge-Decoupled Functionally Invariant Path With Synthetic Personal Data for Personalized ASRabstractFine-tuning generic ASR models with large-scale synthetic personal data can enhance the personalization of ASR models, but it introduces challenges in adapting to synthetic personal data without forgetting real knowledge, and in adapting to personal data without forgetting generic knowledge. Considering that the functionally invariant path (FIP) framework enables model adaptation while preserving prior knowledge, in this letter, we introduce FIP into synthetic-data-augmented personalized ASR models. However, the model still struggles to balance the learning of synthetic, personalized, and generic knowledge when applying FIP to train the model on all three types of data simultaneously. To decouple this learning process and further address the above two challenges, we integrate a gated parameter-isolation strategy into FIP and propose a knowledge-decoupled functionally invariant path (KDFIP) framework, which stores generic and personalized knowledge in separate modules and applies FIP to them sequentially. Specifically, KDFIP adapts the personalized module to synthetic and real personal data and the generic module to generic data. Both modules are updated along personalization-invariant paths, and their outputs are dynamically fused through a gating mechanism. With augmented synthetic data, KDFIP achieves a 29.38% relative character error rate reduction on target speakers and maintains comparable generalization performance to the unadapted ASR baseline. Zhihao Du, Ying Shi 0001, Jiqing Han 0001, Yongjun He 0002 |
IEEE Signal Process. Lett. | 3 |
| 2024 | Serialized Output Training by Learned Dominance
Ying Shi 0001, Lantian Li, Dong Wang 0013, Jiqing Han 0001 |
INTERSPEECH | 1 |
| 2024 | Few-Shot Keyword Spotting from Mixed SpeechabstractFew-shot keyword spotting (KWS) aims to detect unknown keywords with limited training samples.A commonly used approach is the pre-training and fine-tuning framework.While effective in clean conditions, this approach struggles with mixed keyword spotting -simultaneously detecting multiple keywords blended in an utterance, which is crucial in real-world applications.Previous research has proposed a Mix-Training (MT) approach to solve the problem, however, it has never been tested in the few-shot scenario.In this paper, we investigate the possibility of using MT and other relevant methods to solve the two practical challenges together: few-shot and mixed speech.Experiments conducted on the LibriSpeech and Google Speech Command corpora demonstrate that MT is highly effective on this task when employed in either the pre-training phase or the fine-tuning phase.Moreover, combining SSL-based large-scale pre-training (HuBert) and MT fine-tuning yields very strong results in all the test conditions. Junming Yuan, Ying Shi 0001, Lantian Li, Dong Wang 0013, Askar Hamdulla |
INTERSPEECH | 2 |
| 2024 | Keyword Guided Target Speech RecognitionabstractThis letter presents a new target speech recognition problem, where the target speech is defined by a keyword. For instance, when a person speaks “Hey Google” or “Help Me”, we hope the model can recognize the entire contextual speech of that person, even with strong interference speech from other people. The new problem is denoted by target content ASR (TC-ASR). The core challenge of TC-ASR is that the model needs to simultaneously detect the existence of the keyword from heavily mixed speech and recognize the target speech component using the information of the detected keyword segment. Surprisingly, our experiments show that an attention encoder-decoder (AED) model augmented with a keyword encoder can solve this problem pretty well. We also defined a key content spotting (KCS) task and tested the proposed model on it. Our experiments on the LibriMix dataset demonstrated that our approach could address the KCS task with a promising accuracy, outperforming two baseline models by a large margin. Further analysis shows that the proposed model identifies the target speech by a timbre cue, i.e., ensuring that the identified speech is coherent in speaker trait. Ying Shi 0001, Lantian Li, Dong Wang 0013, Jiqing Han 0001 |
IEEE Signal Process. Lett. | 1 |
| 2023 | Spot Keywords From Very Noisy and Mixed Speech
Ying Shi 0001, Dong Wang 0013, Lantian Li, Jiqing Han 0001 |
INTERSPEECH | 1 |
| 2021 | Can We Trust Deep Speech Prior?abstractRecently, speech enhancement (SE) based on deep speech prior has attracted much attention, such as the variational auto-encoder with non-negative matrix factorization (VAE-NMF) architecture. Compared to conventional approaches that represent clean speech by shallow models such as Gaussians with a low-rank covariance, the new approach employs deep generative models to represent the clean speech, which often provides a better prior. Despite the clear advantage in theory, we argue that deep priors must be used with much caution, since the likelihood produced by a deep generative model does not always coincide with the speech quality. We designed a comprehensive study on this issue and demonstrated that based on deep speech priors, a reasonable SE performance can be achieved, but the results might be suboptimal. A careful analysis showed that this problem is deeply rooted in the disharmony between the flexibility of deep generative models and the nature of the maximum-likelihood (ML) training. Ying Shi 0001, Zhiyuan Tang, Lantian Li, Dong Wang 0013, Jiqing Han 0001 |
SLT | 1 |
| 2019 | Gaussian-constrained Training for Speaker VerificationabstractNeural models, in particular the d-vector and x-vector architectures, have produced state-of-the-art performance on many speaker verification tasks. However, two potential problems of these neural models deserve more investigation. Firstly, both models suffer from `information leak', which means that some parameters participating in model training will be discarded during inference, i.e, the layers that are used as the classifier. Secondly, these models do not regulate the distribution of the derived speaker vectors. This `unconstrained distribution' may degrade the performance of the subsequent scoring component, e.g., PLDA. This paper proposes a Gaussian-constrained training approach that (1) discards the parametric classifier, and (2) enforces the distribution of the derived speaker vectors to be Gaussian. Our experiments on the VoxCeleb and SITW databases demonstrated that this new training approach produced more representative and regular speaker embeddings, leading to consistent performance improvement. Lantian Li, Zhiyuan Tang, Ying Shi 0001, Dong Wang 0013 |
ICASSP | 3 |
| 2018 | Deep Factorization for Speech SignalabstractVarious informative factors mixed in speech signals, leading to great difficulty when decoding any of the factors. An intuitive idea is to factorize each speech frame into individual informative factors, though it turns out to be highly difficult. Recently, we found that speaker traits, which were assumed to be long-term distributional properties, are actually short-time patterns, and can be learned by a carefully designed deep neural network (DNN). This discovery motivated a cascade deep factorization (CDF) framework that will be presented in this paper. The proposed framework infers speech factors in a sequential way, where factors previously inferred are used as conditional variables when inferring other factors. We will show that this approach can effectively factorize speech signals, and using these factors, the original speech spectrum can be recovered with a high accuracy. This factorization and reconstruction approach provides potential values for many speech processing tasks, e.g., speaker recognition and emotion recognition, as will be demonstrated in the paper. Lantian Li, Dong Wang 0013, Yixiang Chen 0003, Ying Shi 0001, Zhiyuan Tang, Thomas Fang Zheng |
ICASSP | 4 |
| 2017 | Memory visualization for gated recurrent neural networks in speech recognitionabstractRecurrent neural networks (RNNs) have shown clear superiority in sequence modeling, particularly the ones with gated units, such as long short-term memory (LSTM) and gated recurrent unit (GRU). However, the dynamic properties behind the remarkable performance remain unclear in many applications, e.g., automatic speech recognition (ASR). This paper employs visualization techniques to study the behavior of LSTM and GRU when performing speech recognition tasks. Our experiments show some interesting patterns in the gated memory, and some of them have inspired simple yet effective modifications on the network structure. We report two of such modifications: (1) lazy cell update in LSTM, and (2) shortcut connections for residual learning. Both modifications lead to more comprehensible and powerful networks. Zhiyuan Tang, Ying Shi 0001, Dong Wang 0013, Yang Feng 0004, Shiyue Zhang 0001 |
ICASSP | 2 |
| 2017 | Deep Speaker Feature Learning for Text-Independent Speaker VerificationabstractRecently deep neural networks (DNNs) have been used to learn speaker features.However, the quality of the learned features is not sufficiently good, so a complex back-end model, either neural or probabilistic, has to be used to address the residual uncertainty when applied to speaker verification, just as with raw features.This paper presents a convolutional timedelay deep neural network structure (CT-DNN) for speaker feature learning.Our experimental results on the Fisher database demonstrated that this CT-DNN can produce highquality speaker features: even with a single feature (0.3 seconds including the context), the EER can be as low as 7.68%.This effectively confirmed that the speaker trait is largely a deterministic short-time property rather than a long-time distributional pattern, and therefore can be extracted from just dozens of frames. Lantian Li, Yixiang Chen 0003, Ying Shi 0001, Zhiyuan Tang, Dong Wang 0013 |
INTERSPEECH | 3 |