EDBT 2026 Demo / reviewers in the wild / expert
Zhenchun Lei
dblp:19/16
· DBLP profile ↗
18ranked-venue papers
10as first author
11since 2021 · last 2025
0000-0001-5846-7402ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 10 first-author · 8 since 2021Artificial intelligence and machine learning · 12 · 8 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ConMSDMamba: Multi-Scale Dilated Mamba Based on Conformer for Speech Emotion RecognitionabstractAlthough the Conformer model excels in speech processing, its core self-attention mechanism is limited in capturing multi-scale temporal dynamics and lacks explicit modeling of frequency-domain features, both crucial for Speech Emotion Recognition (SER). To address this, we propose ConMSDMamba, a novel Conformer-based architecture for SER. Specifically, to overcome the single-scale limitation of the original self-attention, we introduce a multi-scale dilated structure with parallel dilated convolutions to capture diverse temporal contexts. We further find that combining this structure with bidirectional Mamba models long-range temporal dependencies more efficiently than multi-head self-attention. Furthermore, to complement the Conformer's time-domain focus, we design a time-frequency convolution module that incorporates a wavelet-based branch for joint time-frequency perception. Experimental results on the widely used IEMOCAP and MELD datasets demonstrate that ConMSDMamba outperforms state-of-the-art methods. Guangyuan Qian, Zhenchun Lei, Sihong Liu, Changhong Liu, Aiwen Jiang |
IEEE Signal Process. Lett. | 2 |
| 2024 | Gmm-Resnext: Combining Generative and Discriminative Models for Speaker VerificationabstractWith the development of deep learning, many different network architectures have been explored in speaker verification. However, most network architectures rely on a single deep learning architecture, and hybrid networks combining different architectures have been little studied in ASV task. In this paper, we propose the GMM-ResNext model for speaker verification. Conventional GMM does not consider the score distribution of each frame feature over all Gaussian components and ignores the relationship between neighboring speech frames. So, we extract the log Gaussian probability features based on the raw acoustic features and use ResNext-based network as the backbone to extract the speaker embedding. GMM-ResNext combines generative and discriminative models to improve the generalization ability of deep learning models and allows one to more easily specify meaningful priors on model parameters. A two-path GMM-ResNext model based on two gender-related GMMs has also been proposed. The experimental results show that the proposed GMM-ResNext achieves relative improvements of 48.1% and 11.3% in EER compared with ResNet34 and ECAPA-TDNN on VoxCeleb1-O. Zhenchun Lei, Changhong Liu |
ICASSP | 2 |
| 2024 | GMM-ResNet2: Ensemble of Group Resnet Networks for Synthetic Speech DetectionabstractDeep learning models are widely used for speaker recognition and spoofing speech detection. We propose the GMM-ResNet2 for synthesis speech detection. Compared with the previous GMM-ResNet model, GMM-ResNet2 has four improvements. Firstly, the different order GMMs have different capabilities to form smooth approximations to the feature distribution, and multiple GMMs are used to extract multi-scale Log Gaussian Probability features. Secondly, the grouping technique is used to improve the classification accuracy by exposing the group cardinality while reducing both the number of parameters and the training time. The final score is obtained by ensemble of all group classifier outputs using the averaging method. Thirdly, the residual block is improved by including one activation function and one batch normalization layer. Finally, an ensemble-aware loss function is proposed to integrate the independent loss functions of all ensemble members. On the ASVspoof 2019 LA task, the GMM-ResNet2 achieves a minimum t-DCF of 0.0227 and an EER of 0.79%. On the ASVspoof 2021 LA task, the GMM-ResNet2 achieves a minimum t-DCF of 0.2362 and an EER of 2.19%, and represents a relative reductions of 31.4% and 76.3% compared with the LFCC-LCNN baseline. Zhenchun Lei, Changhong Liu, Minglei Ma |
ICASSP | 1 |
| 2024 | Music-driven Character Dance Video Generation based on Pre-trained Diffusion ModelabstractLarge-scale pre-trained models have shown significant progress in cross-modal generation tasks, especially in the text-to-image generation task. However, the pre-trained models for audio-guided video are rare. ControlNet [1] provides a new architecture to enhance the pre-trained diffusion models with task-specific conditions. Following the ControlNet [1] architecture, we propose a music-driven character dance video generation model based on the pre-trained diffusion model by taking the text prompt, music, and character image as the additional guidance conditions to generate dance videos. In this model, multimodal semantic correspondence between text, music, and video is exploited to generate character dance videos better by incorporating the pre-trained CLIP [2] and Wav2CLIP [3] models. Additionally, we design a text prompt to improve the appearance quality of the generated character images. Extensive experiments on the AIST++ dataset show the effectiveness of our method and its ability to generate character dance videos effectively. Changhong Liu, Juan Cai, Ji Ye, Zhenchun Lei, Aiwen Jiang |
IJCNN | 5 |
| 2024 | MMIDM: Generating 3D Gesture from Multimodal Inputs with Diffusion Models
Ji Ye, Changhong Liu, Haocong Wan, Aiwen Jiang, Zhenchun Lei |
PRCV (6) | 5 |
| 2023 | Snippet-level Supervised Contrastive Learning-based Transformer for Temporal Action DetectionabstractAnchor-free temporal action detection methods have recently achieved many good results in solving the problem of flexible boundaries and different duration of actions. But the anchor-free methods use local features to predict the action boundaries so that it is sensitive to noises and prone to generate incomplete action proposals. Moreover, there exist long-term temporal dependencies between actions and temporal semantic consistency between action primitives in the same classes of actions. Therefore, we propose a snippet-level supervised contrastive learning-based transformer (SSCL-T) model for temporal action detection, which can learn semantically local and global temporal relationships in actions. This model learns the local temporal dynamic features of actions through local temporal coding and uses the transformer to model the global semantic dependencies between long-term actions. In addition, we utilize the action class information to learn the high-level semantic features of actions by designing a snippet-level supervised contrastive learning, forcing the temporal dynamic features of the same class of actions to be as close as possible and the features of different classes of actions to be as far away as possible, thus effectively realizing accurate prediction of action boundaries. Our model has been verified on two benchmark datasets ActivityNet-v1.3 and THUMOS14. The experimental results demonstrate that the proposed model has significantly improved on both datasets. Compared with the benchmark method BMN, the average mAP value has increased by 2.91% and 8.4% on ActivityNet-v1.3 and THUMOS14, respectively. Ronghai Xu, Changhong Liu, Zhenchun Lei |
IJCNN | 4 |
| 2023 | Group GMM-ResNet for Detection of Synthetic Speech Attacks
Zhenchun Lei, Yingen Yang, Changhong Liu, Minglei Ma |
INTERSPEECH | 1 |
| 2022 | Two-Path GMM-ResNet and GMM-SENet for ASV Spoofing DetectionabstractThe automatic speaker verification system is sometimes vulnerable to various spoofing attacks. The 2-class Gaussian Mixture Model classifier for genuine and spoofed speech is usually used as the baseline for spoofing detection. However, the GMM classifier does not separately consider the scores of feature frames on each Gaussian component. In addition, the GMM accumulates the scores on all frames independently, and does not consider their correlations. We propose the two-path GMM-ResNet and GMM-SENet models for spoofing detection, whose input is the Gaussian probability features based on two GMMs trained on genuine and spoofed speech respectively. The models consider not only the score distribution on GMM components, but also the relationship between adjacent frames. A two-step training scheme is applied to improve the system robustness. Experiments on the ASVspoof 2019 show that the LFCC+GMM-ResNet system can relatively reduce min-tDCF and EER by 76.1% and 76.3% on logical access scenario compared with the GMM, and the LFCC+GMM-SENet system by 94.4% and 95.4% on physical access scenario. After score fusion, the systems give the second-best results on both scenarios. Zhenchun Lei, Hui Yang 0007, Changhong Liu, Minglei Ma, Yingen Yang |
ICASSP | 1 |
| 2022 | Multi-Scale Cascaded Generator for Music-driven Dance SynthesisabstractDance is a creative performance art and must keep coherent with the rhythm and style of music. To address these issues, most of the existing music-driven dance synthesis methods utilize deep generative models and capture the dynamic characteristics of dance motions. However, we observe that dance motions contain big-scale body part movements and small-scale joint movements that are mutually coordinated and related, and the generated dance motions, particularly affected by$L_{1}$loss, are too restrictive and conservative. In this paper, we propose a multi-scale cascaded music-driven dance synthesis network (MC-MDSN) that first generates big-scale body motions conditioned on music and then further refines local small-scale joint motions. Furthermore, we design a multi-scale feature loss to capture the dynamic characteristics of each scale motions and the relations between different scale motion joints. Experimental results show that our method generates better dance motions than the baselines. Changhong Liu, Aiwen Jiang, Zhenchun Lei, Mingwen Wang 0001 |
IJCNN | 5 |
| 2022 | Multi-Path GMM-MobileNet Based on Attack Algorithms and Codecs for Synthetic Speech and Deepfake Detection
Zhenchun Lei, Yingen Yang, Changhong Liu, Minglei Ma |
INTERSPEECH | 2 |
| 2022 | Music-to-Dance Generation with Multiple ConformerabstractIt is necessary for the music-to-dance generation to consider both the kinematics in dance that is highly complex and non-linear and the connection between music and dance movement that is far from deterministic. Existing approaches attempt to address the limited creativity problem, but it is still a very challenging task. First, it is a long-term sequence-to-sequence task. Second, it is noisy in the extracted motion keypoints. Last, there exist local and global dependencies in the music sequence and the dance motion sequence. To address these issues, we propose a novel autoregressive generative framework that predicts future motions based on past motions and music. This framework contains a music conformer, a motion conformer, and a cross-modal conformer, which utilizes the conformer to encode music and motion sequences, and further adapt the cross-modal conformer to the noisy dance motion data that enable it to not only capture local and global dependencies among the sequences but also reduce the effect of noisy data. Quantitative and qualitative experimental results on the publicly available music-to-dance dataset demonstrate our method improves greatly upon the baselines and can generate long-term coherent dance motions well-coordinated with the music. Mingao Zhang, Changhong Liu, Zhenchun Lei, Mingwen Wang 0001 |
ICMR | 4 |
| 2020 | Siamese Convolutional Neural Network Using Gaussian Probability Feature for Spoofing Speech Detection
Zhenchun Lei, Yingen Yang, Changhong Liu, Jihua Ye |
INTERSPEECH | 1 |
| 2016 | Mahalanobis Metric Scoring Learned from Weighted Pairwise Constraints in I-Vector Speaker Recognition System
Zhenchun Lei, Yanhong Wan, Yingen Yang |
INTERSPEECH | 1 |
| 2011 | Maximum Likelihood i-vector Space Using PCA for Speaker Verification
Zhenchun Lei, Yingchun Yang |
INTERSPEECH | 1 |
| 2010 | Combining the Likelihood and the Kullback-Leibler Distance in Estimating the Universal Background Model for Speaker Verification Using SVMabstractThe state-of-the-art methods for speaker verification are based on the support vector machine. The Gaussian supervector SVM is a typical method which uses the Gaussian mixture model for creating “feature vectors” for the discriminative SVM. And all GMMs are adapted from the same universal background model, which is got by maximum likelihood estimation on a large number of data sets. So the UBM should cover the feature space widely as possible. We propose a new method to estimate the parameters of the UBM by combining the likelihood and the Kullback-Leibler distances in the UBM. Its aim is to find the model parameters which get the high likelihood value and all Gaussian distributions are dispersed to cover the feature space in a great measuring. Experiments on NIST 2001 task show that our method can improve the performance obviously. Zhenchun Lei |
ICPR | 1 |
| 2009 | UBM-based sequence kernel for speaker recognition
Zhenchun Lei |
INTERSPEECH | 1 |
| 2006 | A discriminative method for speaker verification using the difference information
Zhenchun Lei, Yingchun Yang, Zhaohui Wu 0001 |
INTERSPEECH | 1 |
| 2005 | Mixture of support vector machines for text-independent speaker recognitionabstractIn this paper, the mixture of support vector machines is proposed and applied to text-independent speaker recognition. The mixture of experts is used and is implemented by the divide-and-conquer approach. The purpose of adopting this idea is to deal with the large scale speech data and improve the performance of speaker recognition. The principle is to train several parallel SVMs on the subsets of the whole dataset and then combine them in the distance or probabilistic fashion. The experiments have been run on the YOHO database, and the results show that the mixture model is superior to the basic Gaussian mixture model. Zhenchun Lei, Yingchun Yang, Zhaohui Wu 0001 |
INTERSPEECH | 1 |