EDBT 2026 Demo / reviewers in the wild / expert
Peng Zhang 0148
dblp:21/1048-148
· DBLP profile ↗
2ranked-venue papers in the field
1as first author
2since 2021 · last 2025
0000-0002-4851-2094ORCID · verified
Domains — venue-derived; a paper can count in several
Other / Interdisciplinary · 2 (1 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Speech Emotion Recognition via Multi-Level Acoustic Modeling and Cross-Modal Temporal FusionabstractSpeech Emotion Recognition (SER) aims to enable machines to understand human affective states using only audio signals. However, effectively extracting and leveraging comprehensive acoustic information remains a challenging task. In this paper, we propose an end-to-end multi-level acoustic modeling framework with a newly designed Cross-Temporal Affective Fusion (CTAF) module for dimensional emotion prediction (e.g., valence, arousal and dominance). Specifically, we first apply a dynamic chunk segmentation strategy, where each speech signal is divided into overlapping chunks using an adaptive step size. The step size is automatically adjusted according to each utterance’s length, ensuring a fixed number of equal-duration chunks. From these chunks, we extract three levels of acoustic features: Low-Level Descriptors (LLDs), Mel-Frequency Cepstral Coefficient (MFCC), and WavLM representations. To model each feature type effectively, we design specialized encoders: a Fusion-Aware Temporal Convolutional Network (FATCNet) for capturing local temporal dynamics in LLDs, a CNN-LSTM architecture for modeling both static and dynamic patterns in MFCCs, and the WavLM for extracting global contextual semantics. The extracted features are then treated as multi-modal inputs and fused using the proposed CTAF module. Experiments on the MSP-Podcast and IEMOCAP datasets demonstrate that our model achieves competitive SER performance. Peng Zhang 0148 |
MMAsia | 2 |
| 2023 | Speech Spoofing Detection Based on Graph Attention Networks with Spectral and Temporal InformationabstractAutomatic speaker verification (ASV) systems are vulnerable to synthetic speech attacks. Synthetic algorithms usually introduce artifacts in specific sub-bands or time segments. However, under unknown spoofing attacks, it is challenging to choose the right domain for effective detection. In this paper, we propose a speech spoofing detection method based on graph attention networks with spectral and temporal information. First, high-level features of raw audio are extracted using SENet channel attention to enhance the spatial correlation between speech frames. Then, spectral graph and temporal graph are constructed for the high-level features using graph attention networks. Finally, we design a new heterogeneous multi-domain co-graph attention module to process the information from different domains for effective speech spoofing detection. The proposed model was evaluated on the ASVspoof 2019 dataset and obtains a min t-DCF of 0.0264 and an EER of 0.94%, exhibiting competitive performance. Experiments also show its effectiveness when detecting unknown types of attacks. Peng Zhang 0148, Meijuan Li |
MMAsia | 1 |