VLDB 2026 Research / reviewers in the wild / expert
Yijie Li 0001
dblp:54/8054-1
· DBLP profile ↗
11ranked-venue papers
0as first author
6since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 6 since 2021Artificial intelligence and machine learning · 5 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Cross-Modal Parallel Training for Improving end-to-end Accented Speech RecognitionabstractMulti-accent speech recognition is a key challenge in current speech recognition due to the pronunciation variations of different accents. In this study, we propose a Cross-modal Parallel Training (CPT) approach for improving the accent robustness of state-of-the-art Conformer-Transducer (Conformer-T) ASR system. Specifically, in CPT, a novel cross-modal attention and fusion module is first designed as a frontend to align the low-level acoustic speech representations with phonetic embeddings, and thus normalizing accent variations into a shared standard pronunciation latent space; Then, a parallel training mechanism is proposed to simultaneously model both the acoustic and accent normalized multi-modal features for improving accented ASR performance. Moreover, different multi-objective training losses with text-induced and phonetic target units are also investigated. Our experiments are performed on the public CommonVoice English accented ASR tasks, results show that the proposed CPT outperforms the strong baseline by relative 9.3%-13.4% WER reductions on six evaluation sets, all without increasing any model parameters or computational costs during ASR inference. Renchang Dong, Yijie Li 0001, Dongxing Xu, Yanhua Long |
ICASSP | 2 |
| 2024 | Accent-Specific Vector Quantization for Joint Unsupervised and Supervised Training in Accent Robust Speech RecognitionabstractHow to effectively use limited supervised accent data to improve the accented ASR is of paramount importance. In this work, we propose an accent-specific quantization for joint unsupervised and supervised training (AQ-JUST) of end-to-end ASR models to address this issue. Specifically, two variants of AQ-JUST are investigated, namely BAQ-JUST and SAQ-JUST, by employing different model structures and training methods to capture the distinctiveness and commonalities between diverse accents, thus enhancing the performance of accented ASR systems. Our experiments are performed on both accented English and Mandarin ASR tasks. Results show that the proposed methods outperform the strong JUST baseline by relative 3.9% to 9.4% word/character error rate reductions on accented test sets. Yijie Li 0001, Dongxing Xu, Yanhua Long |
ICASSP | 2 |
| 2023 | Advanced RawNet2 with Attention-based Channel Masking for Synthetic Speech Detection
Yanhua Long, Yijie Li 0001, Dongxing Xu |
INTERSPEECH | 3 |
| 2023 | Multi-pass Training and Cross-information Fusion for Low-resource End-to-end Accented Speech Recognition
Yanhua Long, Yijie Li 0001 |
INTERSPEECH | 3 |
| 2022 | Selective Pseudo-labeling and Class-wise Discriminative Fusion for Sound Event DetectionabstractIn recent years, exploring effective sound separation (SSep) techniques to improve overlapping sound event detection (SED) attracts more and more attention.Creating accurate separation signals to avoid the catastrophic error accumulation during SED model training is very important and challenging.In this study, we first propose a novel selective pseudo-labeling approach, termed SPL, to produce high confidence separated target events from blind sound separation outputs.These target events are then used to fine-tune the original SED model that pre-trained on the sound mixtures in a multi-objective learning style.Then, to further leverage the SSep outputs, a class-wise discriminative fusion is proposed to improve the final SED performances, by combining multiple frame-level event predictions of both sound mixtures and their separated signals.All experiments are performed on the public DCASE 2021 Task 4 dataset, and results show that our approaches significantly outperforms the official baseline, the collar-based F 1, PSDS1 and PSDS2 performances are improved from 44.3%, 37.3% and 54.9% to 46.5%, 44.5% and 75.4%, respectively. Yunhao Liang, Yanhua Long, Yijie Li 0001, Jiaen Liang |
INTERSPEECH | 3 |
| 2021 | Multi-Channel Target Speech Extraction with Channel Decorrelation and Target Speaker AdaptationabstractThe end-to-end approaches for single-channel target speech extraction have attracted widespread attention. However, the studies for end-to-end multi-channel target speech extraction are still relatively limited. In this work, we propose two methods for exploiting the multi-channel spatial information to extract the target speech. The first one is using a target speech adaptation layer in a parallel encoder architecture. The second one is designing a channel decorrelation mechanism to extract the inter-channel differential information to enhance the multi-channel encoder representation. We compare the proposed methods with two strong state-of-the-art baselines. Experimental results on the multi-channel reverberant WSJ0 2-mix dataset demonstrate that our proposed methods achieve up to 11.2% and 11.5% relative improvements in SDR and SiSDR respectively, which are the best reported results on this task to the best of our knowledge. Jiangyu Han, Xinyuan Zhou, Yanhua Long, Yijie Li 0001 |
ICASSP | 4 |
| 2020 | Multi-Encoder-Decoder Transformer for Code-Switching Speech RecognitionabstractCode-switching (CS) occurs when a speaker alternates words of two or more languages within a single sentence or across sentences.Automatic speech recognition (ASR) of CS speech has to deal with two or more languages at the same time.In this study, we propose a Transformer-based architecture with two symmetric language-specific encoders to capture the individual language attributes, that improve the acoustic representation of each language.These representations are combined using a language-specific multi-head attention mechanism in the decoder module.Each encoder and its corresponding attention module in the decoder are pre-trained using a large monolingual corpus aiming to alleviate the impact of limited CS training data.We call such a network a multi-encoder-decoder (MED) architecture.Experiments on the SEAME corpus show that the proposed MED architecture achieves 10.2% and 10.8% relative error rate reduction on the CS evaluation sets with Mandarin and English as the matrix language respectively. Xinyuan Zhou, Emre Yilmaz 0001, Yanhua Long, Yijie Li 0001, Haizhou Li 0001 |
INTERSPEECH | 4 |
| 2018 | Active Learning for LF-MMI Trained Neural Networks in ASR
Yanhua Long, Yijie Li 0001, Jiaen Liang |
INTERSPEECH | 3 |
| 2018 | Offline to online speaker adaptation for real-time deep neural network based LVCSR systems
Yanhua Long, Yijie Li 0001 |
Multim. Tools Appl. | 2 |
| 2009 | iFLY system for the NIST 2008 speaker recognition evaluationabstractThe description of iFLY system submitted for NIST 2008 speaker recognition evaluation (SRE), which has achieved excellent performance in the 2008 SRE evaluation, is presented in this paper. Our primary system is a fusion of two subsystems GMM-UBM and GMM-SVM. For each sub-system, two kinds of short-time acoustic features PLP and LPCC are adopted. We focus on three key issues in this evaluation: channel compensation, multi-lingual or bi-lingual cues and the voice activity detection. We also point out that data selection and factor analysis play key roles in the system improvement. Wu Guo, Yanhua Long, Yijie Li 0001, Eryu Wang, Li-Rong Dai 0001 |
ICASSP | 3 |
| 2009 | The I4U system in NIST 2008 speaker recognition evaluationabstractThis paper describes the performance of the I4U speaker recognition system in the NIST 2008 Speaker Recognition Evaluation. The system consists of seven subsystems, each with different cepstral features and classifiers. We describe the I4U Primary system and report on its core test results as they were submitted, which were among the best-performing submissions. The I4U effort was led by the Institute for Infocomm Research, Singapore (IIR), with contributions from the University of Science and Technology of China (USTC), the University of New South Wales, Australia (UNSW), Nanyang Technological University, Singapore (NTU) and Carnegie Mellon University, USA (CMU). Haizhou Li 0001, Bin Ma 0001, Kong-Aik Lee, Hanwu Sun, Donglai Zhu, Khe Chai Sim, Chang Huai You, Rong Tong, Ismo Kärkkäinen, Chien-Lin Huang, Vladimir Pervouchine, Wu Guo, Yijie Li 0001, Li-Rong Dai 0001, Mohaddeseh Nosratighods, Tharmarajah Thiruvaran, Julien Epps, Eliathamby Ambikairajah, Chng Eng Siong, Tanja Schultz, Qin Jin |
ICASSP | 13 |