VLDB 2026 Research / reviewers in the wild / expert
Ming-Hui Wu
dblp:288/1657
· DBLP profile ↗
8ranked-venue papers
0as first author
8since 2021 · last 2024
0009-0002-0179-1441ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer graphics and multimedia
1 paper |
Audio and music processing · 75% Multimedia analysis and retrieval · 25% |
Topics — the 4 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Audio and music processing › speech recognition
acoustic modeling |
0.7 | 1 | 2023 | A Semi-Supervised Complementary Joint Training Approach for Low-Resource Speech Recognition · IEEE ACM Trans. Audio Speech Lang. Process. 2023 |
Audio and music processing › speech recognition
low-resource speech recognition |
0.7 | 1 | 2023 | A Semi-Supervised Complementary Joint Training Approach for Low-Resource Speech Recognition · IEEE ACM Trans. Audio Speech Lang. Process. 2023 |
Multimedia analysis and retrieval
semi-supervised learning |
0.7 | 1 | 2023 | A Semi-Supervised Complementary Joint Training Approach for Low-Resource Speech Recognition · IEEE ACM Trans. Audio Speech Lang. Process. 2023 |
Audio and music processing
speech recognition |
0.7 | 1 | 2023 | A Semi-Supervised Complementary Joint Training Approach for Low-Resource Speech Recognition · IEEE ACM Trans. Audio Speech Lang. Process. 2023 |
Methods — techniques the papers use, named apart from their topics
text-to-speech synthesis · 0.7pseudo-labeling · 0.7iterative training · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | SLPA-Net: A Real-Time Recognition Network for Intelligent Stomata Localization and Phenotypic AnalysisabstractPlant stomatal phenotype traits play an important role in improving crop water use efficiency, stress resistance and yield. However, at present, the acquisition of phenotype traits mainly relies on manual measurement, which is time-consuming and laborious. In order to obtain high-throughput stomatal phenotype traits, we proposed a real-time recognition network SLPA-Net for stomata localization and phenotypic analysis. After locating and identifying stomatal density data, ellipse fitting is used to automatically obtain phenotype data such as apertures. Aiming at the problems of small stomata and high similarity to background, we introduced ECANet to improve the accuracy of stoma and aperture location. In order to effectively alleviate the unbalance problem in bounding box regression, we replaced the Loss function with a more effective Focal EIoU Loss. The experimental results show that SLPA-Net has excellent performance in the migration generalization and robustness of stomata and apertures detection and identification, as well as the correlation between stomata phenotype data obtained and artificial data. Ye-Tong Wang, Ming-Hui Wu, Cheng-Long Zhou, Chen Zheng 0002, Si-Yi Guo, Chun-Peng Song |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2023 | Monotonic Gaussian regularization of attention for robust automatic speech recognition
Ye-Qian Du, Ming-Hui Wu, Zhouwang Yang |
Comput. Speech Lang. | 2 |
| 2023 | A Semi-Supervised Complementary Joint Training Approach for Low-Resource Speech RecognitionabstractBoth unpaired speech and text have shown to be beneficial for low-resource automatic speech recognition (ASR), which, however were either separately used for pre-training, self-training and language model (LM) training, or jointly used for designing hybrid models in literature. In this work, we leverage both unpaired speech and text to train a general ASR model, which are used in the form of data pairs by generating the missing parts in prior to model training. We propose to train a model alternatively using the prepared speech-PseudoLabel and SynthesizedAudio-text pairs and reveal the complementary property in both acoustic and linguistic features. The proposed method is thus called complementary joint training (CJT). Based on the basic CJT, label masking for pseudo-labels and parallel layers for synthesized audio are then proposed for re-training to further cope with the deviations from real data, termed as CJT++. In addition, the proposed CJT is extended to the scenario with zero paired data by considering an iterative CJT for the training of seed ASR model. Experimental results on Libri-light show the efficacy of joint training as well as two second-round training strategies, and the superiority over recent models is validated, particularly in extreme low-resource cases. Ye-Qian Du, Jie Zhang 0042, Ming-Hui Wu, Zhouwang Yang |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2022 | A Noise-Robust Self-Supervised Pre-Training Model Based Speech Representation Learning for Automatic Speech RecognitionabstractWav2vec2.0 is a popular self-supervised pre-training framework for learning speech representations in the context of automatic speech recognition (ASR). It was shown that wav2vec2.0 has a good robustness against the domain shift, while the noise robustness is still unclear. In this work, we therefore first analyze the noise robustness of wav2vec2.0 via experiments. We observe that wav2vec2.0 pre-trained on noisy data can obtain good representations and thus improve the ASR performance on the noisy test set, which however brings a performance degradation on the clean test set. To avoid this issue, in this work we propose an enhanced wav2vec2.0 model. Specifically, the noisy speech and the corresponding clean version are fed into the same feature encoder, where the clean speech provides training targets for the model. Experimental results reveal that the proposed method can not only improve the ASR performance on the noisy test set which surpasses the original wav2vec2.0, but also ensure a tiny performance decrease on the clean test set. In addition, the effectiveness of the proposed method is demonstrated under different types of noise conditions. Qiushi Zhu, Jie Zhang 0042, Ming-Hui Wu, Li-Rong Dai 0001 |
ICASSP | 4 |
| 2022 | Learning Contextually Fused Audio-Visual Representations For Audio-Visual Speech RecognitionabstractWith the advance in self-supervised learning for audio and visual modalities, it has become possible to learn a robust audio-visual speech representation. This would be beneficial for improving the audio-visual speech recognition (AVSR) performance, as the multi-modal inputs contain more fruitful information in principle. In this paper, based on existing self-supervised representation learning methods for audio modality, we therefore propose an audio-visual representation learning approach. The proposed approach explores both the complementarity of audio-visual modalities and long-term context dependency using a transformer-based fusion module and a flexible masking strategy. After pre-training, the model is able to extract fused representations required by AVSR. Without loss of generality, it can be applied to single-modal tasks, e.g., audio/visual speech recognition by simply masking out one modality in the fusion module. The proposed pre-trained model is evaluated on speech recognition and lipreading tasks using one or two modalities, where the superiority is revealed. Jie Zhang 0042, Jianshu Zhang 0001, Ming-Hui Wu, Li-Rong Dai 0001 |
ICIP | 4 |
| 2022 | An Experimental Comparison between Low-Resource Semi-Supervised and High-Resource Supervised Automatic Speech Recognition ModelsabstractAutomatic speech recognition (ASR) is an important module in many multimedia applications. Recently, semi-supervised ASR has attracted increasing attention, which can be classified into self-supervised learning and self-training. It was shown that the combination of the two methods is beneficial for ASR on open datasets, e.g., LibriSpeech. However, it is still not completely clear how the semi-supervised model behaves on more challenging industrial datasets compared with supervised approaches using high-resource labeled data. In this paper, we therefore present an experimental study on the combination of self-supervised learning (e.g., wav2vec 2.0) and self-training (e.g., noisy student) in a low-resource industrial setting. The in-domain pre-trained and fine-tuned wav2vec 2.0 model is utilized to teach a baseline VGG-Transformer by pseudo-labeling in an industrial setting. Results reveal that without extra language models the combined semi-supervised acoustic model using much less labeled data can perform as well as the supervised counterpart, which would be rather beneficial for the application of semi-supervised ASR models in low-resource scenarios. Ao-Ran Gan, Jie Zhang 0042, Ming-Hui Wu, Li-Rong Dai 0001 |
ICME | 3 |
| 2022 | A Complementary Joint Training Approach Using Unpaired Speech and Text A Complementary Joint Training Approach Using Unpaired Speech and Text
Ye-Qian Du, Jie Zhang 0042, Qiushi Zhu, Li-Rong Dai 0001, Ming-Hui Wu, Zhouwang Yang |
INTERSPEECH | 5 |
| 2021 | An Improved Wav2Vec 2.0 Pre-Training Approach Using Enhanced Local Dependency Modeling for Speech Recognition
Qiushi Zhu, Jie Zhang 0042, Ming-Hui Wu, Li-Rong Dai 0001 |
Interspeech | 3 |