Hongyang Chen 0004

dblp:13/3715-4 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
6since 2021 · last 2025
0000-0002-7626-0162ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2025 PGD-N2L: A Parameter-Guided Disentanglement Approach for Normal-To-Lombard Speech Conversion
abstract
The Normal-To-Lombard (N2L) speech conversion can effectively improve speech intelligibility in noisy communication scenarios and serve as a data augmentation tool for various speech-related algorithms. However, existing N2L methods did not aim to disentangle the Lombard effect from other speech attributes, leading to incomplete conversions. In this paper, we propose a Parameter-Guided Disentanglement approach for N2L speech conversion (PGD-N2L) which decomposes speech into linguistic content, speaker identity, and Lombard effect. To extract disentangled linguistic content, we propose a DeLomb-Based content encoder. To extract disentangled speaker identity and Lombard effect, we propose a style encoder that combines a fine-tuned speaker encoder and a learnable Lombard encoder to form a personalized style embedding. Furthermore, an En-Lomb-Based injection module is designed to accurately integrate the target Lombard effect and speaker identity into the linguistic content based on personalized style embedding, ensuring complete Lombard conversion. Experimental results demonstrate that our proposed method outperforms existing N2L models in speech intelligibility, acoustic similarity, and speech quality. Ablation studies confirm that the fine-tuned speaker encoder and the De-Lomb block effectively improve speech intelligibility and acoustic similarity, while the En-Lomb block enables the converted speech to more closely match the target Lombard speech.
Hongyang Chen 0004, Yuhong Yang 0001, Xinmeng Xu, Weiping Tu, Zhongyuan Wang 0001, Cedar Lin
ICME1
2024 EMALG: An Enhanced Mandarin Lombard Grid Corpus with Meaningful Sentences
abstract
This study investigates the Lombard effect, where individuals adapt their speech in noisy environments. We introduce an enhanced Mandarin Lombard grid (EMALG) corpus with meaningful sentences, enhancing the Mandarin Lombard grid (MALG) corpus. EMALG features 34 speakers and improves recording setups, addressing challenges faced by MALG with nonsense sentences. Our findings reveal that in Mandarin, meaningful sentences are more effective in enhancing the Lombard effect. Additionally, we uncover that female exhibit a more pronounced Lombard effect than male when uttering meaningful sentences. Moreover, our results reaffirm the consistency in the Lombard effect comparison between English and Mandarin found in previous research.
Baifeng Li, Qingmu Liu, Yuhong Yang 0001, Hongyang Chen 0004, Weiping Tu
ICASSP4
2024 Exploring Sentence Type Effects on the Lombard Effect and Intelligibility Enhancement: A Comparative Study of Natural and Grid Sentences
Hongyang Chen 0004, Yuhong Yang 0001, Zhongyuan Wang 0001, Weiping Tu, Haojun Ai, Cedar Lin
INTERSPEECH1
2023 PMMSD: Development of the Matrix Sentence Intelligibility Dataset for Mandarin with Lombard Effect
abstract
This paper presents a Paired Mandarin Matrix Sentence Dataset (PMMSD), which will be available after publication. PMMSD is the first Mandarin matrix sentence intelligibility dataset containing both plain and Lombard speech for scientific research. The results verify that different Lombard styles would affect word intelligibility to different degrees and the Lombard effect helps maintain homogeneous intelligibility against contextual interference. All of the discoveries indicate that the Lombard effect should be considered when building intelligibility datasets with noise in the future.
Hanchen Pei, Yuhong Yang 0001, Xufeng Chen, Qingmu Liu, Hongyang Chen 0004, Weiping Tu
ICASSP5
2022 CS-CTCSCONV1D: Small footprint speaker verification with channel split time-channel-time separable 1-dimensional convolution
Linjun Cai, Yuhong Yang 0001, Xufeng Chen, Weiping Tu, Hongyang Chen 0004
INTERSPEECH5
2022 Mandarin Lombard Grid: a Lombard-grid-like corpus of Standard Chinese
Yuhong Yang 0001, Xufeng Chen, Qingmu Liu, Weiping Tu, Hongyang Chen 0004, Linjun Cai
INTERSPEECH5