VLDB 2026 Research / reviewers in the wild / expert
Shenghui Lu
dblp:383/9643
· DBLP profile ↗
7ranked-venue papers
2as first author
7since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CCGS: A Cross-modal Collaborative Gradient Sparsification for Accelerating Distributed Multimodal Model Training
Shenghui Lu, Waixi Liu 0001, Jinhuang Huang, Qingchun Chen |
Euro-Par (2) | 1 |
| 2025 | Dynamic Language Group-based MoE: Enhancing Code-Switching Speech Recognition with Hierarchical RoutingabstractThe Mixture of Experts (MoE) model is a promising approach for handling code-switching speech recognition (CS-ASR) tasks. However, the existing CS-ASR work on MoE has yet to leverage the advantages of MoE’s parameter scaling ability fully. This work proposes DLG-MoE, a Dynamic Language Group-based MoE, which can effectively handle the CS-ASR task and leverage the advantages of parameter scaling. DLG-MoE operates based on a hierarchical routing mechanism. First, the language router explicitly models the language attribute and dispatches the representations to the corresponding language expert groups. Subsequently, the unsupervised router within each language group implicitly models attributes beyond language and coordinates expert routing and collaboration. DLG-MoE outperforms the existing MoE methods on CS-ASR tasks while demonstrating great flexibility. It supports different top-k inference and streaming capabilities and can also prune the model parameters flexibly to obtain a monolingual sub-model. Hukai Huang, Shenghui Lu, Yahui Shan, He Qu, Fengrun Zhang, Wenhao Guan, Qingyang Hong |
ICASSP | 2 |
| 2025 | SlimSpeech: Lightweight and Efficient Text-to-Speech with Slim Rectified FlowabstractRecently, flow matching based speech synthesis has significantly enhanced the quality of synthesized speech while reducing the number of inference steps. In this paper, we introduce SlimSpeech, a lightweight and efficient speech synthesis system based on rectified flow. We have built upon the existing speech synthesis method utilizing the rectified flow model, modifying its structure to reduce parameters and serve as a teacher model. By refining the reflow operation, we directly derive a smaller model with a more straight sampling trajectory from the larger model, while utilizing distillation techniques to further enhance the model performance. Experimental results demonstrate that our proposed method, with significantly reduced model parameters, achieves comparable performance to larger models through one-step sampling. Kaidi Wang 0001, Wenhao Guan, Shenghui Lu, Jianglong Yao, Qingyang Hong |
ICASSP | 3 |
| 2025 | A Noisy Label Filter based on GMM Binary Classification for Speaker VerificationabstractNoisy labels are inevitable in real-world datasets. These noisy labels cause deep neural networks to gradient descent towards the wrong direction, leading to performance degradation. In this paper, We propose an efficient method for filtering out noisy labels during training. We calculate an embedding center for each speaker and compute the cosine similarity between each audio embedding and its corresponding speaker embedding center. Subsequently, we use a Gaussian Mixture Model (GMM) to classify the samples into two categories based on their cosine similarities and filter out the class with lower similarity as noisy labels. These steps are iterated until training is complete. We conducted experiments on real-world datasets as well as scenarios with artificially added noisy labels of different types. The experimental results demonstrate that our method achieved notable performance. Jianglong Yao, Shenghui Lu, Qinyang Hong |
ICASSP | 2 |
| 2025 | A Two-Stage Hierarchical Deep Filtering Framework for Real-Time Speech Enhancement
Shenghui Lu, Hukai Huang, Jinanglong Yao, Qingyang Hong |
INTERSPEECH | 1 |
| 2025 | Pseudo Labels-based Neural Speech Enhancement for the AVSR Task in the MISP-Meeting Challenge
Longjie Luo, Shenghui Lu, Qingyang Hong |
INTERSPEECH | 2 |
| 2024 | MinSpeech: A Corpus of Southern Min Dialect for Automatic Speech Recognition
Jiayan Lin, Shenghui Lu, Hukai Huang, Wenhao Guan, Hui Bu, Qingyang Hong |
INTERSPEECH | 2 |