Ye-Qian Du

dblp:317/7047 · DBLP profile ↗
← Back
4ranked-venue papers
3as first author
4since 2021 · last 2025
0000-0002-6176-8676ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
1 paper
Audio and music processing · 75% Multimedia analysis and retrieval · 25%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Audio and music processing › speech recognition
acoustic modeling
0.712023
A Semi-Supervised Complementary Joint Training Approach for Low-Resource Speech Recognition · IEEE ACM Trans. Audio Speech Lang. Process. 2023
Audio and music processing › speech recognition
low-resource speech recognition
0.712023
A Semi-Supervised Complementary Joint Training Approach for Low-Resource Speech Recognition · IEEE ACM Trans. Audio Speech Lang. Process. 2023
Multimedia analysis and retrieval
semi-supervised learning
0.712023
A Semi-Supervised Complementary Joint Training Approach for Low-Resource Speech Recognition · IEEE ACM Trans. Audio Speech Lang. Process. 2023
Audio and music processing
speech recognition
0.712023
A Semi-Supervised Complementary Joint Training Approach for Low-Resource Speech Recognition · IEEE ACM Trans. Audio Speech Lang. Process. 2023

Methods — techniques the papers use, named apart from their topics

text-to-speech synthesis · 0.7pseudo-labeling · 0.7iterative training · 0.7
YearPublicationVenuePosition
2025 Robust Multimodal Representation under Uncertain Missing Modalities
abstract
Multimodal representation learning has gained significant attention across various fields, yet it faces challenges when dealing with missing modalities in real-world applications. Existing solutions are confined to specific scenarios, such as single-modality missing or missing modalities in test cases, thereby restricting their applicability. To address a more general scenario of uncertain missing modalities in both training and testing phases, we propose Robust Multimodal Representation under Uncertain Missing Modalities (RMRU). This framework projects each modality’s representation into a shared subspace, enabling the reconstruction of any missing modalities within a unified model. We propose an interaction refinement module that utilizes cross-modal attention to enhance these reconstructions, particularly beneficial in scenarios with limited complete modality data. Furthermore, we introduce an iterative training strategy that alternately trains different modules to effectively utilize both complete and incomplete modality data. Experimental results on four benchmark datasets demonstrate the superiority of RMRU over existing baselines, particularly in scenarios with a high rate of missing modalities. Remarkably, our proposed RMRU can be broadly applied to diverse scenarios, regardless of modality types and quantities.
Guilin Lan, Ye-Qian Du, Zhouwang Yang
ACM Trans. Multim. Comput. Commun. Appl.2
2023 Monotonic Gaussian regularization of attention for robust automatic speech recognition
Ye-Qian Du, Ming-Hui Wu, Zhouwang Yang
Comput. Speech Lang.1
2023 A Semi-Supervised Complementary Joint Training Approach for Low-Resource Speech Recognition
abstract
Both unpaired speech and text have shown to be beneficial for low-resource automatic speech recognition (ASR), which, however were either separately used for pre-training, self-training and language model (LM) training, or jointly used for designing hybrid models in literature. In this work, we leverage both unpaired speech and text to train a general ASR model, which are used in the form of data pairs by generating the missing parts in prior to model training. We propose to train a model alternatively using the prepared speech-PseudoLabel and SynthesizedAudio-text pairs and reveal the complementary property in both acoustic and linguistic features. The proposed method is thus called complementary joint training (CJT). Based on the basic CJT, label masking for pseudo-labels and parallel layers for synthesized audio are then proposed for re-training to further cope with the deviations from real data, termed as CJT++. In addition, the proposed CJT is extended to the scenario with zero paired data by considering an iterative CJT for the training of seed ASR model. Experimental results on Libri-light show the efficacy of joint training as well as two second-round training strategies, and the superiority over recent models is validated, particularly in extreme low-resource cases.
Ye-Qian Du, Jie Zhang 0042, Ming-Hui Wu, Zhouwang Yang
IEEE ACM Trans. Audio Speech Lang. Process.1
2022 A Complementary Joint Training Approach Using Unpaired Speech and Text A Complementary Joint Training Approach Using Unpaired Speech and Text
Ye-Qian Du, Jie Zhang 0042, Qiushi Zhu, Li-Rong Dai 0001, Ming-Hui Wu, Zhouwang Yang
INTERSPEECH1