Yizhou Peng

dblp:277/1404 · DBLP profile ↗
← Back
14ranked-venue papers
2as first author
14since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 1 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Evaluating the Expressive Appropriateness of Speech in Rich Contexts
abstract
Tianrui Wang, Ziyang Ma, Yizhou Peng, Haoyu Wang, Zhikang Niu, Zikang Huang, Yihao Wu, Yi-Wen Chao, Yu Jiang, Yuheng Lu, Guanrou Yang, Xuanchen Li, Hexin Liu, Chunyu Qiang, Cheng Gong, Yifan Yang, Tianchi Liu, Junyu Wang, Nana Hou, Meng Ge, Fuming You, Yang Wei, Zhongqian Sun, Hu Haifeng, Xiaobao Wang, Eng Siong Chng, Xie Chen, Longbiao Wang, Jianwu Dang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Tianrui Wang, Ziyang Ma 0001, Yizhou Peng, Zhikang Niu, Zikang Huang, Yi-Wen Chao, Yuheng Lu, Guanrou Yang, Xuanchen Li, Hexin Liu, Chunyu Qiang, Yifan Yang 0005, Tianchi Liu 0004, Nana Hou, Meng Ge, Fuming You, Zhongqian Sun, Haifeng Hu 0009, Xiaobao Wang, Chng Eng Siong, Xie Chen 0001, Longbiao Wang, Jianwu Dang 0001
ACL (1)3
2026 Spatiotemporal-Attention-Based Channel Prediction for UAV-RIS-Assisted LEO Satellite MIMO Communications
abstract
Low Earth orbit (LEO) satellite communications play a critical role in achieving global connectivity, yet they face significant challenges due to high satellite mobility and incomplete channel state information (CSI). Moreover, the integration of reconfigurable intelligent surfaces (RIS) in certain scenarios introduces additional complexities. In this paper, we propose a novel MIMO channel prediction framework tailored for LEO satellite communications involving unmanned aerial vehicle-mounted RIS (UAV-RIS), employing a spatiotemporal-attention (ST-attention) mechanism to capture both the spatial correlations among antennas and the temporal dynamics of rapidly varying channels. Furthermore, we leverage masked pretraining to enhance the model’s robustness under scenarios of severe CSI incompleteness, enabling effective reconstruction of missing channel information. Comprehensive simulations demonstrate that our approach outperforms traditional model-based predictors, whether historical CSI is fully available or only partially observed.
Yizhou Peng, Ruofei Ma, Gongliang Liu, Weixiao Meng 0001, Carla Fabiana Chiasserini, Roberto Garello
IEEE Trans. Wirel. Commun.2
2025 A-SMiLE: Affective Sparse Mixture-of-Experts Adapter with Multi-Task Learning for Spoken Dialogue Models
Yi-Wen Chao, Yizhou Peng, Dianwen Ng, Chongjia Ni, Bin Ma 0001, Chng Eng Siong
INTERSPEECH2
2025 FD-Bench: A Full-Duplex Benchmarking Pipeline Designed for Full Duplex Spoken Dialogue Systems
Yizhou Peng, Yi-Wen Chao, Dianwen Ng, Chongjia Ni, Bin Ma 0001, Chng Eng Siong
INTERSPEECH1
2025 Language-Aware Prompt Tuning for Parameter-Efficient Seamless Language Expansion in Multilingual ASR
abstract
Recent advancements in multilingual automatic speech recognition (ASR) have been driven by large-scale end-to-end models like Whisper. However, challenges such as language interference and expanding to unseen languages (language expansion) without degrading performance persist. This paper addresses these with three contributions: 1) Entire Soft Prompt Tuning (Entire SPT), which applies soft prompts to both the encoder and decoder, enhancing feature extraction and decoding; 2) Language-Aware Prompt Tuning (LAPT), which leverages cross-lingual similarities to encode shared and language-specific features using lightweight prompt matrices; 3) SPT-Whisper, a toolkit that integrates SPT into Whisper and enables efficient continual learning. Experiments across three languages from FLEURS demonstrate that Entire SPT and LAPT outperform Decoder SPT by 5.0% and 16.0% in language expansion tasks, respectively, providing an efficient solution for dynamic, multilingual ASR models with minimal computational overhead.
Sheng Li 0010, Hao Huang 0009, Ayiduosi Tuohan, Yizhou Peng
INTERSPEECH5
2025 Adapting Whisper for Parameter-efficient Code-Switching Speech Recognition via Soft Prompt Tuning
abstract
Large-scale multilingual ASR models like Whisper excel in high-resource settings but face challenges in low-resource scenarios, such as rare languages and code-switching (CS), due to computational costs and catastrophic forgetting. We explore Soft Prompt Tuning (SPT), a parameter-efficient method to enhance CS ASR while preserving prior knowledge. We evaluate two strategies: (1) full fine-tuning (FFT) of both soft prompts and the entire Whisper model, demonstrating improved cross-lingual capabilities compared to traditional methods, and (2) adhering to SPT's original design by freezing model parameters and only training soft prompts. Additionally, we introduce SPT4ASR, a combination of different SPT variants. Experiments on the SEAME and ASRU2019 datasets show that deep prompt tuning is the most effective SPT approach, and our SPT4ASR methods achieve further error reductions in CS ASR, maintaining parameter efficiency similar to LoRA, without degrading performance on existing languages.
Yizhou Peng, Hao Huang 0009, Sheng Li 0010
INTERSPEECH2
2024 ED-sKWS: Early-Decision Spiking Neural Networks for Rapid, and Energy-Efficient Keyword Spotting
Zeyang Song, Qianhui Liu, Qu Yang, Yizhou Peng, Haizhou Li 0001
INTERSPEECH4
2024 Near-field channel estimation for extremely large-scale Terahertz communications
Songjie Yang, Yizhou Peng, Wanting Lyu, Hongjun He, Zhongpei Zhang, Chau Yuen
Sci. China Inf. Sci.2
2023 Adapting Code-Switching Language Models with Statistical-Based Text Augmentation
Chaiyasait Prachaseree, Thi-Nga Ho, Yizhou Peng, Kyaw Zin Tun, Chng Eng Siong, G. S. S. Chalapthi
ACIIDS (2)4
2023 Self-supervised Learning Representation based Accent Recognition with Persistent Accent Memory
Zhiwei Xie 0007, Haihua Xu 0001, Yizhou Peng, Hexin Liu, Hao Huang 0009, Chng Eng Siong
INTERSPEECH4
2022 Minimum Word Error Training For Non-Autoregressive Transformer-Based Code-Switching ASR
abstract
Non-autoregressive end-to-end ASR framework might be potentially appropriate for code-switching recognition task thanks to its inherent property that present output token being independent of historical ones. However, it still under-performs the state-of-the-art autoregressive ASR frameworks. In this paper, we propose various approaches to boosting the performance of a CTC-mask-based non-autoregressive Transformer under code-switching ASR scenario. To begin with, we attempt diversified masking method that are closely related with code-switching point, yielding an improved baseline model. More importantly, we employ Minimum Word Error (MWE) criterion to train the model. One of the challenges is how to generate a diversified hypothetical space, so as to obtain the average loss for a given ground truth. To address such a challenge, we explore different approaches to yielding desired N-best-based hypothetical space. We demonstrate the efficacy of the proposed methods on SEAME corpus, a challenging English-Mandarin code-switching corpus for Southeast Asia community. Compared with the cross-entropy-trained strong baseline, the proposed MWE training method achieves consistent performance improvement on the test sets.
Yizhou Peng, Haihua Xu 0001, Hao Huang 0009, Chng Eng Siong
ICASSP1
2022 Mining Hard Samples Locally And Globally For Improved Speech Separation
abstract
Speech separation dataset typically consists of hard and non-hard samples, and the former is minority and latter majority. The data imbalance problem biases the model towards non-hard samples and weakens the generalization capability. Given that the average separation performance is sufficiently good, improving hard samples may contribute more to back-end tasks. In this paper, we propose two methods to alleviate data imbalance in speech separation task, based on local and global hard sample mining. For the local, we propose weighted loss to compensate for hard samples by increasing their weights in each batch. For the global, we perform global hard sample mining and re-sample to increase the proportion of hard samples in the training set. Because hard sample mining using objective loss in dynamic mixing leads to local results, we propose an indirect method using speaker-specific parameters, based on the fact that pitch median difference and x-vector cosine distance of two speakers in a mixture are closely correlated with separation SI-SNRi. Experimental results show that both methods decrease the percentage of hard samples in the test set than using dynamic mixing only while keeping the average SI-SNRi comparable, and the global method shows more promising results than the local one.
Yizhou Peng, Hao Huang 0009, Ying Hu 0005, Sheng Li 0010
ICASSP2
2022 Multi-Semantic Path Representation Learning for Travel Time Estimation
abstract
Travel time estimation of a given path is a crucial task of Intelligent Transportation Systems (ITS). Accurate travel time estimation can benefit multiple downstream applications such as route planning, real-time navigation, and urban construction. However, it is a challenging problem since the travel time is largely affected by multiple complicated factors including spatial factors, temporal factors and external factors, and obtaining informative representations of a given path is not trivial. Most previous works solved this problem in either Euclidean space or non-Euclidean space, which was unilateral to represent the actual traveling path and led to relatively poor performance. To address this, this paper proposes a multi-semantic path representation method to exploit information in Euclidean space and non-Euclidean space simultaneously. First, since the path is composed of several segments, we generate semantic representations of segments in non-Euclidean space by taking both the time information and the historical co-occurrence into consideration. Second, as the path could be equally represented as several travelled intersections, semantic representations of intersection sequences are also extracted to improve the capability of the method by considering information in Euclidean space. Meanwhile, semantic representations from properties, including the length and the type of segments, are also incorporated into the model. Finally, a sequence learning component is added on the top to aggregate the information along the entire path and provides the final estimation. Extensive experiments were conducted on two real-world taxi trajectories datasets, and the experimental results demonstrate the superiority of the proposed method.
Liangzhe Han, Bowen Du 0001, Jingjing Lin, Leilei Sun, Xucheng Li, Yizhou Peng
IEEE Trans. Intell. Transp. Syst.6
2021 E2E-Based Multi-Task Learning Approach to Joint Speech and Accent Recognition
abstract
In this paper, we propose a single multi-task learning framework to perform End-to-End (E2E) speech recognition (ASR) and accent recognition (AR) simultaneously.The proposed framework is not only more compact but can also yield comparable or even better results than standalone systems.Specifically, we found that the overall performance is predominantly determined by the ASR task, and the E2E-based ASR pretraining is essential to achieve improved performance, particularly for the AR task.Additionally, we conduct several analyses of the proposed method.First, though the objective loss for the AR task is much smaller compared with its counterpart of ASR task, a smaller weighting factor with the AR task in the joint objective function is necessary to yield better results for each task.Second, we found that sharing only a few layers of the encoder yields better AR results than sharing the overall encoder.Experimentally, the proposed method produces WER results close to the best standalone E2E ASR ones, while it achieves 7.7% and 4.2% relative improvement over standalone and single-task-based joint recognition methods on test set for accent recognition respectively.
Yizhou Peng, Van Tung Pham, Haihua Xu 0001, Hao Huang 0009, Chng Eng Siong
Interspeech2