VLDB 2026 Research / reviewers in the wild / expert
Guangcun Wei
dblp:275/2341
· DBLP profile ↗
8ranked-venue papers
1as first author
8since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GAGE: Generative Adversarial Enhancement for High-Fidelity Audio Watermark Removal
Guangcun Wei, Chengde Zhang, Yanhong Long |
KSEM (4) | 2 |
| 2026 | SE-EEND: A Structurally Enhanced End-to-End Neural Diarization System
Penghao Ma, Guangcun Wei, Chuike Kong, Jianfeng Fang |
MMM (2) | 2 |
| 2025 | MixSENet: A Lightweight Model for Speech Enhancement with Multi-Scale Features and Contextual ModelingabstractIn this paper, we propose a novel lightweight speech enhancement network, MixSENet (Mixed Structure Speech Enhancement Network), designed to address the challenges associated with multi-scale feature processing of speech signals and coping with complex noise environments. The network is based on the U-Net architecture and innovatively introduces a hybrid structural block, which consists of a multi-scale parallel large convolutional kernel module (MSPLCK) and an enhanced parallel attention module (EPAM). The MSPLCK achieves large sensory fields and multi-scale feature extraction through parallel dilated convolution, whereas the EPAM can simultaneously process global shared information and local time-frequency features effectively coping with uneven noise distributions. Extensive experiments demonstrate that MixSENet shows significant performance advantages on the Voice Bank + DEMAND dataset, with a parameter count of only 0.43M, significantly reducing training and deployment costs. The method has substantial practical value in real application scenarios, especially for resource-constrained mobile devices and embedded systems. Chuike Kong, Guangcun Wei, Penghao Ma |
ICMR | 2 |
| 2024 | Dual-Stage Training Frame-level Overlapping Speech Detection Model Based on Convolutional Neural Network ArchitectureabstractA dual-stage training frame-level overlapping speech detection model is proposed in this paper, with its core being based on RepVGG blocks structure, and RepVGG is built on simple one-dimensional Convolutional Neural Networks (CNNs). Both the learning of local features and temporal features of acoustic signals are accomplished by the RepVGG structure. The model categorizes speech into triple classes: non-speech frames, non-overlapping speech frames, and overlapping speech frames. Firstly, RepVGG blocks are employed in the construction of a Speech Activity Detection (SAD) model to detect non-speech and speech frames. Subsequently, a Overlapping Speech Detection (OSD) model is constructed using RepVGG blocks and is trained with real labels to differentiate between non-overlapping and overlapping speech frames. Finally, these two models are combined and fine-tuned to directly output classification results. The proposed overlapping speech detection model achieves advanced performance on the AliMeeting dataset with fewer parameters. Zhifei Pan, Guangcun Wei, Qingge Fang, Jihua Fu |
IJCNN | 2 |
| 2024 | DocPointer: A parameter-efficient Pointer Network for Key Information Extraction
Guangcun Wei, Haochen Xu, Boyan Guo |
MMAsia | 2 |
| 2024 | MRGAN: LightWeight Monaural Speech Enhancement Using GAN Network
Chunyu Meng, Guangcun Wei, Yanhong Long, Chuike Kong, Penghao Ma |
PRCV (4) | 2 |
| 2024 | Spoofing Speech Detection Method Based on Self-supervised Front End and Feature Enhancement
Boyan Guo, Guangcun Wei, Chunyu Meng, Chengde Zhang |
PRICAI (4) | 2 |
| 2023 | End-to-end speaker identification research based on multi-scale SincNet and CGAN
Guangcun Wei, Hang Min |
Neural Comput. Appl. | 1 |