Guangcun Wei

dblp:275/2341 · DBLP profile ↗
← Back
8ranked-venue papers
1as first author
8since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021
YearPublicationVenuePosition
2026 GAGE: Generative Adversarial Enhancement for High-Fidelity Audio Watermark Removal
Guangcun Wei, Chengde Zhang, Yanhong Long
KSEM (4)2
2026 SE-EEND: A Structurally Enhanced End-to-End Neural Diarization System
Penghao Ma, Guangcun Wei, Chuike Kong, Jianfeng Fang
MMM (2)2
2025 MixSENet: A Lightweight Model for Speech Enhancement with Multi-Scale Features and Contextual Modeling
abstract
In this paper, we propose a novel lightweight speech enhancement network, MixSENet (Mixed Structure Speech Enhancement Network), designed to address the challenges associated with multi-scale feature processing of speech signals and coping with complex noise environments. The network is based on the U-Net architecture and innovatively introduces a hybrid structural block, which consists of a multi-scale parallel large convolutional kernel module (MSPLCK) and an enhanced parallel attention module (EPAM). The MSPLCK achieves large sensory fields and multi-scale feature extraction through parallel dilated convolution, whereas the EPAM can simultaneously process global shared information and local time-frequency features effectively coping with uneven noise distributions. Extensive experiments demonstrate that MixSENet shows significant performance advantages on the Voice Bank + DEMAND dataset, with a parameter count of only 0.43M, significantly reducing training and deployment costs. The method has substantial practical value in real application scenarios, especially for resource-constrained mobile devices and embedded systems.
Chuike Kong, Guangcun Wei, Penghao Ma
ICMR2
2024 Dual-Stage Training Frame-level Overlapping Speech Detection Model Based on Convolutional Neural Network Architecture
abstract
A dual-stage training frame-level overlapping speech detection model is proposed in this paper, with its core being based on RepVGG blocks structure, and RepVGG is built on simple one-dimensional Convolutional Neural Networks (CNNs). Both the learning of local features and temporal features of acoustic signals are accomplished by the RepVGG structure. The model categorizes speech into triple classes: non-speech frames, non-overlapping speech frames, and overlapping speech frames. Firstly, RepVGG blocks are employed in the construction of a Speech Activity Detection (SAD) model to detect non-speech and speech frames. Subsequently, a Overlapping Speech Detection (OSD) model is constructed using RepVGG blocks and is trained with real labels to differentiate between non-overlapping and overlapping speech frames. Finally, these two models are combined and fine-tuned to directly output classification results. The proposed overlapping speech detection model achieves advanced performance on the AliMeeting dataset with fewer parameters.
Zhifei Pan, Guangcun Wei, Qingge Fang, Jihua Fu
IJCNN2
2024 DocPointer: A parameter-efficient Pointer Network for Key Information Extraction
Guangcun Wei, Haochen Xu, Boyan Guo
MMAsia2
2024 MRGAN: LightWeight Monaural Speech Enhancement Using GAN Network
Chunyu Meng, Guangcun Wei, Yanhong Long, Chuike Kong, Penghao Ma
PRCV (4)2
2024 Spoofing Speech Detection Method Based on Self-supervised Front End and Feature Enhancement
Boyan Guo, Guangcun Wei, Chunyu Meng, Chengde Zhang
PRICAI (4)2
2023 End-to-end speaker identification research based on multi-scale SincNet and CGAN
Guangcun Wei, Hang Min
Neural Comput. Appl.1