Youjun Chen

dblp:21/3305 · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
7since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 5 since 2021Theory of computation · 2 · 2 first-author
YearPublicationVenuePosition
2025 Phone-purity Guided Discrete Tokens for Dysarthric Speech Recognition
abstract
Discrete tokens provide compact and domain-adaptable representations of speech features. However, their application to disordered speech, characterized by articulation imprecision and significant mismatch with normal voice, remains unexplored. To this end, this paper proposes novel phone-purity guided (PPG) discrete tokens to address the weakened phonetic discrimination arising during unsupervised K-means clustering or vector quantization of continuous features. Phonetic label supervision is incorporated to regularize the maximum likelihood and reconstruction error costs in standard K-means and VAE-VQ-based token extraction. Experiments on the UASpeech corpus show that PPG-based discrete tokens extracted from HuBERT consistently outperform hybrid TDNN and End-to-End (E2E) Conformer systems using non-PPG tokens. Statistically significant word error rate (WER) reductions of up to 0.99% and 1.77% absolute (3.21% and 4.82% relative) are achieved across varying codebook sizes for the 16 UASpeech test dysarthric speakers. The lowest WER of 23.25% is obtained by combining systems using complementary token features. Consistent improvements are also observed in phone purity, and t-SNE visualizations demonstrate sharper decision boundaries between K-means/VAE-VQ clusters with the introduction of phone-purity guidance.
Huimeng Wang, Xurong Xie, Mengzhe Geng, Shujie Hu, Haoning Xu, Youjun Chen, Zhaoqing Li, Jiajun Deng, Xunying Liu
ICASSP6
2025 Effective and Efficient Mixed Precision Quantization of Speech Foundation Models
abstract
This paper presents a novel mixed-precision quantization approach for speech foundation models that tightly integrates mixed-precision learning and quantized model parameter estimation into one single model compression stage. Experiments conducted on LibriSpeech dataset with fine-tuned wav2vec2.0-base and HuBERT-large models suggest the resulting mixed-precision quantized models increased the lossless compression ratio by factors up to 1.7x and 1.9x over the respective uniform-precision and two-stage mixed-precision quantized baselines that perform precision learning and model parameters quantization in separate and disjointed stages, while incurring no statistically word error rate (WER) increase over the 32-bit full-precision models. The system compression time of wav2vec2.0-base and HuBERT-large models is reduced by up to 1.9 and 1.5 times over the two-stage mixed-precision baselines, while both produce lower WERs. The best-performing 3.5-bit mixed-precision quantized HuBERT-large model produces a lossless compression ratio of 8.6x over the 32-bit full-precision system.
Haoning Xu, Zhaoqing Li, Zengrui Jin, Huimeng Wang, Youjun Chen, Guinan Li, Mengzhe Geng, Shujie Hu, Jiajun Deng, Xunying Liu
ICASSP5
2025 Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition
Youjun Chen, Xurong Xie, Haoning Xu, Mengzhe Geng, Guinan Li, Chengxi Deng, Huimeng Wang, Shujie Hu, Xunying Liu
INTERSPEECH1
2025 MOPSA: Mixture of Prompt-Experts Based Speaker Adaptation for Elderly Speech Recognition
Chengxi Deng, Xurong Xie, Shujie Hu, Mengzhe Geng, Yicong Jiang, Jiankun Zhao, Jiajun Deng, Guinan Li, Youjun Chen, Huimeng Wang, Haoning Xu, Xunying Liu
INTERSPEECH9
2025 Towards One-bit ASR: Extremely Low-bit Conformer Quantization Using Co-training and Stochastic Precision
Zhaoqing Li, Haoning Xu, Zengrui Jin, Lingwei Meng, Tianzi Wang, Huimeng Wang, Youjun Chen, Shujie Hu, Xunying Liu
INTERSPEECH7
2025 Effective and Efficient One-pass Compression of Speech Foundation Models Using Sparsity-aware Self-pinching Gates
Haoning Xu, Zhaoqing Li, Youjun Chen, Huimeng Wang, Guinan Li, Mengzhe Geng, Chengxi Deng, Xunying Liu
INTERSPEECH3
2024 Joint Speaker Features Learning for Audio-visual Multichannel Speech Separation and Recognition
abstract
Interspeech 2024, 1-5 September 2024, Kos, Greece
Guinan Li, Jiajun Deng, Youjun Chen, Mengzhe Geng, Shujie Hu, Zhe Li 0030, Zengrui Jin, Tianzi Wang, Xurong Xie, Helen M. Meng, Xunying Liu
INTERSPEECH3
2016 Two-stage scheduling on identical machines with assignable delivery times to minimize the maximum delivery completion time
abstract
In this paper, we consider the two-stage scheduling problem in which n jobs are first processed on m identical machines at a manufacturing facility and then delivered to their customers by one vehicle which can deliver one job at each shipment. In the problem, a set of n delivery times is given in advance, and in a schedule, the n delivery times should be assigned to the n jobs, respectively. The objective is to minimize the maximum delivery completion time, i.e., the time when all jobs are delivered to their respective customers and the vehicle returns to the facility. For this problem, we present a 32-approximation algorithm and a polynomial-time approximation scheme.
Youjun Chen, Lingfa Lu, Jinjiang Yuan
Theor. Comput. Sci.1
2015 Preemptive scheduling on identical machines with delivery coordination to minimize the maximum delivery completion time
Youjun Chen, Lingfa Lu, Jinjiang Yuan
Theor. Comput. Sci.1
2008 By using grey area relational grade combined with NLP method to optimize GM(1, 1) model
abstract
In this paper, we suggest a new optimization method by using grey area relational grade combined with nonlinear programming (abbreviated to NLP) method for GM(1,1) model’s parameters, we introduce the grey area relational grade to establish a NLP optimization parameters model for a GM(1,1) model that has been established, by using mathematical software LINGO 10.0 for its optimal solution. By a lot of data’s analysis, we conclude the new method is effective and feasible for the GM(1,1) models that have been established by using any other methods.
Youjun Chen, Hongying He, Yong Wei 0001
FUZZ-IEEE1