EDBT 2026 Demo / reviewers in the wild / expert
Shengyu Peng
dblp:405/2191
· DBLP profile ↗
9ranked-venue papers
3as first author
9since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 9 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Recursive Feature Learning from Pre-Trained Models for Spoofing Speech DetectionabstractIt was recently revealed that using features extracted from pre-trained models can achieve much better performance than using conventional hand-crafted acoustic features for spoofing speech detection. In this paper, we therefore enhance the features from pre-trained model based on recursive learning. Specifically, we modify the pre-trained model by feeding the features from the topmost transformer layer to bottom layers recursively, and the obtained recursive features from the bottom layers are fused with that from topmost layer. The fused features are then fed into the backend classifiers. Experiments are carried out on two benchmark datasets (i.e., ASVspoof 2019 LA and ASVspoof 2021 LA), which show the superiority of the proposed method over state-of-the-art systems. Yang Ai, Zuoliang Li, Shengyu Peng, Wu Guo |
ICASSP | 4 |
| 2025 | Aligning Noisy-Clean Speech Pairs at Feature and Embedding Levels for Learning Noise-Invariant Speaker RepresentationsabstractIn this paper, we propose a noise-invariant speaker representation learning (SRL) approach by aligning noisy-clean speech pairs at both the feature and embedding levels for model training. Specifically, we first construct noisy-clean pairs using data augmentation during training. The noisy features are then processed by a Conformer-based enhancement module. The feature-level alignment is achieved by minimizing the mean squared error between the enhanced and original clean data. At the embedding level, we introduce a supervised contrastive learning loss with noise-adaptive margin to simultaneously enhance the intra-speaker compactness and the inter-speaker separability and better adapt different noise levels, in combination with the Barlow Twins self-supervised loss to align the noisy-clean data pairs and reduce noise redundancy in the embedding space. Finally, these loss components are integrated with conventional classification loss to train the SRL network. Experimental results on various VoxCeleb1 test sets synthesized with noise sources demonstrate the effectiveness of the proposed method. Zuoliang Li, Yang Ai, Jie Zhang 0042, Shengyu Peng, Bin Gu 0004, Wu Guo |
ICASSP | 4 |
| 2025 | A Study of Multi-Scale Feature Learning From Pre-Trained Models on Speaker VerificationabstractIn this paper, a multi-scale feature fusion paradigm is proposed to fully exploit the power of the pre-trained models for text-independent speaker verification. It contains a front-end feature extractor and an enhanced ECAPA-TDNN backend in a cascade manner. The feature extractor incorporates local representations of the CNN layers as well as the global clues of the Transformer layers of the pre-trained models, which are combined to construct the multi-scale discriminative features. The outputs of the feature extractor are then fed into the back-end model (tailored from ECAPA-TDNN) to obtain the final speaker embedding. Results on VoxCeleb datasets validate the superiority of the proposed method with equal error rates of 0.633% and 0.457% on the official trials of Vox1-O using the base and large pre-trained models, respectively. Shengyu Peng, Wu Guo, Jie Zhang 0042, Zuoliang Li, Bin Gu 0004, Yang Ai |
ICASSP | 1 |
| 2025 | Leveraging Multi-Level Features of ATST with Conformer-Based Dual-Branch Network for Sound Event Detection
Lipeng Dai, Shengyu Peng, Wu Guo |
INTERSPEECH | 4 |
| 2025 | Parameter-Efficient Fine-tuning with Instance-Aware Prompt and Parallel Adapters for Speaker Verification
Shengyu Peng, Wu Guo, Jie Zhang 0042, Lipeng Dai, Zuoliang Li |
INTERSPEECH | 1 |
| 2024 | Contrastive Learning and Inter-Speaker Distribution Alignment Based Unsupervised Domain Adaptation for Robust Speaker Verification
Zuoliang Li, Wu Guo, Shengyu Peng |
INTERSPEECH | 4 |
| 2024 | Fine-tune Pre-Trained Models with Multi-Level Feature Fusion for Speaker Verification
Shengyu Peng, Wu Guo, Zuoliang Li |
INTERSPEECH | 1 |
| 2024 | Adapter Learning from Pre-trained Model for Robust Spoof Speech Detection
Wu Guo, Shengyu Peng, Zhuhai Li |
INTERSPEECH | 3 |
| 2024 | Spoofing Speech Detection by Modeling Local Spectro-Temporal and Long-term Dependency
Wu Guo, Shengyu Peng |
INTERSPEECH | 5 |