EDBT 2026 Demo / reviewers in the wild / expert
Xuefei Liu
dblp:10/283
· DBLP profile ↗
25ranked-venue papers
3as first author
20since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 17 · 1 first-author · 15 since 2021Artificial intelligence and machine learning · 14 · 1 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Security and privacy · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PSA-MF: Personality-Sentiment Aligned Multi-Level Fusion for Multimodal Sentiment AnalysisabstractMultimodal sentiment analysis (MSA) is a research field that recognizes human sentiments by combining textual, visual, and audio modalities. The main challenge lies in integrating sentiment-related information from different modalities, which typically arises during the unimodal feature extraction phase and the multimodal feature fusion phase. Existing methods extract only shallow information from unimodal features during the extraction phase, neglecting sentimental differences across different personalities. During the fusion phase, they directly merge the feature information from each modality without considering differences at the feature level. This ultimately affects the model's recognition performance. To address this problem, we propose a personality-sentiment aligned multi-level fusion framework. We introduce personality traits during the feature extraction phase and propose a novel personality-sentiment alignment method to obtain personalized sentiment embeddings from the textual modality for the first time. In the fusion phase, we introduce a novel multi-level fusion method. This method gradually integrates sentimental information from textual, visual, and audio modalities through multimodal pre-fusion and a multi-level enhanced fusion strategy. Our method has been evaluated through multiple experiments on two commonly used datasets, achieving state-of-the-art results. Kang Zhu, Zhengqi Wen, Jianhua Tao 0001, Xuefei Liu, Ruibo Fu |
AAAI | 5 |
| 2026 | Personality-aware Multimodal Deception Detection with multimodal large language model
Cong Cai, Zhengqi Wen, Xuefei Liu, Jianhua Tao 0001, Bin Liu 0041 |
Pattern Recognit. | 3 |
| 2026 | CMDPAD: A Chinese multimodal dynamic personality and affect dataset for affect prediction in conversations
Zisen Zhou, Chang Wen, Xuefei Liu, Jianhua Tao 0001, Zhengqi Wen, Zheng Lian 0004, Jinming Zhao, Bingsen Xiong, Shaozheng Qin |
Pattern Recognit. | 4 |
| 2025 | Mel-Refine: A Plug-and-Play Approach to Refine Mel-Spectrogram in Audio GenerationabstractMainstream Text-to-Audio (TTA) models that rely on Mel-spectrograms often struggle to generate audio with rich content, leading to blurred or incoherent outputs. This stems from an inability to model intricate spectral details and textures. We investigate the role of U-Net components in generation and find that high-frequency components in skip-connections and the backbone are crucial for texture, while low-frequency backbone components are vital for the denoising process. Based on this, we propose “Mel-Refine,” a plug-and-play approach that enhances Mel-spectrogram quality by adjusting component weights during inference. Our method requires no additional training or fine-tuning and is fully compatible with any diffusion-based TTA architecture. Experiments show that Mel-Refine boosts the performance of the latest TTA model, Tango2, by $25 \%$, demonstrating its effectiveness. Hongming Guo, Ruibo Fu, Yizhong Geng, Shuchen Shi, Tao Wang 0074, Chunyu Qiang, Ya Li 0001, Zhengqi Wen, Xuefei Liu, Chenxing Li |
ASRU | 10 |
| 2025 | DPI-TTS: Directional Patch Interaction for Fast-Converging and Style Temporal Modeling in Text-to-SpeechabstractIn recent years, speech diffusion models have advanced rapidly. Alongside the widely used U-Net architecture, transformer-based models such as the Diffusion Transformer (DiT) have also gained attention. However, current DiT speech models treat Mel spectrograms as general images, which overlooks the specific acoustic properties of speech. To address these limitations, we propose a method called Directional Patch Interaction for Text-to-Speech (DPI-TTS), which builds on DiT and achieves fast training without compromising accuracy. Notably, DPI-TTS employs a low-to-high frequency, frame-by-frame progressive inference approach that aligns more closely with acoustic properties, enhancing the naturalness of the generated speech. Additionally, we introduce a fine-grained style temporal modeling method that further improves speaker style similarity. Experimental results demonstrate that our method increases the training speed by nearly 2 times and significantly outperforms the baseline models. Ruibo Fu, Zhengqi Wen, Tao Wang 0074, Chunyu Qiang, Jianhua Tao 0001, Chenxing Li, Shuchen Shi, Yuankun Xie, Xuefei Liu, Guanjun Li |
ICASSP | 14 |
| 2025 | Mixture of Experts Fusion for Fake Audio Detection Using Frozen wav2vec 2.0abstractSpeech synthesis technology has posed a serious threat to speaker verification systems. Currently, the most effective fake audio detection methods utilize pretrained models, and integrating features from various layers of pretrained model further enhances detection performance. However, most of the previously proposed fusion methods require fine-tuning the pretrained models, resulting in excessively long training times and hindering model iteration when facing new speech synthesis technology. To address this issue, this paper proposes a feature fusion method based on the Mixture of Experts, which extracts and integrates features relevant to fake audio detection from layer features, guided by a gating network based on the last layer feature, while freezing the pretrained model. Experiments conducted on the ASVspoof2019 and ASVspoof2021 datasets demonstrate that the proposed method achieves competitive performance compared to those requiring fine-tuning. Ruibo Fu, Zhengqi Wen, Jianhua Tao 0001, Yuankun Xie, Shuchen Shi, Chenxing Li, Xuefei Liu, Guanjun Li |
ICASSP | 12 |
| 2025 | MTPareto: A MultiModal Targeted Pareto Framework for Fake News DetectionabstractMultimodal fake news detection is essential for maintaining the authenticity of Internet multimedia information. Significant differences in form and content of multimodal information lead to intensified optimization conflicts, hindering effective model training as well as reducing the effectiveness of existing fusion methods for bimodal. To address this problem, we propose the MTPareto framework to optimize multimodal fusion, using a Targeted Pareto(TPareto) optimization algorithm for fusion-level-specific objective learning with a certain focus. Based on the designed hierarchical fusion network, the algorithm defines three fusion levels with corresponding losses and implements all-modal-oriented Pareto gradient integration for each. This approach accomplishes superior multimodal fusion by utilizing the information obtained from intermediate fusion to provide positive effects to the entire process. Experiment results on FakeSV and FVC datasets show that the proposed framework outperforms baselines and the TPareto optimization algorithm achieves 2.40% and 1.89% accuracy improvement respectively. Kaiying Yan, Moyang Liu, Ruibo Fu, Zhengqi Wen, Jianhua Tao 0001, Xuefei Liu, Guanjun Li |
ICASSP | 7 |
| 2025 | MER 2025: When Affective Computing Meets Large Language ModelsabstractMER2025 is the third year of our MER series of challenges. Previously, MER2023 (http://merchallenge.cn/mer2023) focused on multi-label learning, noise robustness, and semi-supervised learning, while MER2024 (https://zeroqiaoba.github.io/MER2024-website) introduced a new track dedicated to open-vocabulary emotion recognition. This year, MER2025 centers on the theme ''When Affective Computing Meets Large Language Models (LLMs)''. We aim to shift the paradigm from traditional categorical frameworks reliant on predefined emotion taxonomies to LLM-driven generative methods, offering innovative solutions for more accurate and reliable emotion understanding. The challenge contains four tracks: MER-SEMI focuses on fixed categorical emotion recognition enhanced by semi-supervised learning; MER-FG explores fine-grained emotions, expanding recognition from basic to nuanced emotional states; MER-DES incorporates multimodal cues (beyond emotion words) into predictions to enhance model interpretability; MER-PR reveals whether emotion prediction results can improve personality recognition performance. For the first three tracks, the baseline code is available at MERTools (https://github.com/zeroQiaoba/MERTools) and datasets can be accessed via Hugging Face (https://huggingface.co/datasets/MERChallenge/MER2025). For the last track, the dataset and baseline code are available on GitHub (https://github.com/cai-cong/MER25_personality). Zheng Lian 0004, Rui Liu 0008, Kele Xu, Bin Liu 0041, Xuefei Liu, Yazhou Zhang 0001, Xin Liu 0012, Yong Li 0032, Zebang Cheng, Haolin Zuo, Ziyang Ma 0001, Xiaojiang Peng, Xie Chen 0001, Ya Li 0001, Erik Cambria, Guoying Zhao 0001, Björn W. Schuller, Jianhua Tao 0001 |
ACM Multimedia | 5 |
| 2025 | MDPE: A Multimodal Deception Dataset with Personality and Emotional CharacteristicsabstractDeception detection has garnered increasing attention in recent years due to the significant growth of digital media and heightened ethical and security concerns. It has been extensively studied using multimodal methods, including video, audio, and text. In addition, individual differences in deception production and detection are believed to play a crucial role. Although some studies have utilized individual information such as personality traits to enhance the performance of deception detection, current systems remain limited, partly due to a lack of sufficient datasets for evaluating performance. To address this issue, we introduce a multimodal deception dataset MDPE. Besides deception features, this dataset also includes individual differences information in personality and emotional expression characteristics. It can explore the impact of individual differences on deception behavior. It comprises over 104 hours of deception and emotional videos from 193 subjects. Furthermore, we conducted numerous experiments to provide valuable insights for future deception detection research. MDPE not only supports deception detection, but also provides conditions for tasks such as personality recognition and emotion recognition, and can even study the relationships between them. We believe that MDPE will become a valuable resource for promoting research in the field of affective computing. Cong Cai, Shan Liang 0007, Xuefei Liu, Kang Zhu, Zhengqi Wen, Jianhua Tao 0001, Jizhou Cui, Zhenhua Cheng, Hanzhe Xu, Ruibo Fu, Bin Liu 0041 |
ACM Multimedia | 3 |
| 2025 | Dynamic neighbourhood particle swarm optimisation algorithm for solving multi-root direct kinematics in coupled parallel mechanisms
Shikun Wen, Yassine Gharbi, Youzhi Xu, Xuefei Liu, Heow Pueh Lee, Linxian Che, Aihong Ji |
Expert Syst. Appl. | 4 |
| 2024 | Dual-View Multimodal Interaction in Multimodal Sentiment AnalysisabstractOutstanding performance in sentiment analysis not only relies on the design of sophisticated fusion methods but also on the crucial step of designing excellent modal interaction methods. To the best of our knowledge, there are few methods addressing the capture of multimodal spatial features. Majority of feature interactions have been primarily focused on temporal aspects, with less attention given to the combined spatiotemporal feature interaction (SFI). In this paper, we design a dual-view multimodal interaction method, named DVMI, primarily consisting of two parts. In the first part, a triangular convolutional module is proposed for ample temporal interaction between modalities, implicit local and global SFI, and capturing global spatial representations. Building upon the foundation laid in the first part, the second part employs an attention mechanism for explicit global SFI. To demonstrate the effectiveness of the DVMI framework,we conduct extensive experiments on three datasets, achieving state-of-the-art experimental results. Kang Zhu, Cunhang Fan, Jianhua Tao 0001, Jun Xue 0001, Xuefei Liu, Zhengqi Wen, Zhao Lv |
ICME | 6 |
| 2024 | Codecfake: An Initial Dataset for Detecting LLM-based Deepfake Audio
Yuankun Xie, Ruibo Fu, Zhengqi Wen, Jianhua Tao 0001, Xuefei Liu, Shuchen Shi |
INTERSPEECH | 8 |
| 2024 | PPPR: Portable Plug-in Prompt Refiner for Text to Audio Generation
Shuchen Shi, Ruibo Fu, Zhengqi Wen, Jianhua Tao 0001, Tao Wang 0074, Chunyu Qiang, Xuefei Liu |
INTERSPEECH | 9 |
| 2024 | Genuine-Focused Learning using Mask AutoEncoder for Generalized Fake Audio Detection
Ruibo Fu, Zhengqi Wen, Yuankun Xie, Jianhua Tao 0001, Xuefei Liu, Shuchen Shi |
INTERSPEECH | 8 |
| 2024 | Generalized Fake Audio Detection via Deep Stable Learning
Ruibo Fu, Zhengqi Wen, Yuankun Xie, Xuefei Liu, Jianhua Tao 0001, Shuchen Shi |
INTERSPEECH | 7 |
| 2022 | MeRIPseqPipe: an integrated analysis pipeline for MeRIP-seq data based on NextflowabstractSUMMARY: MeRIPseqPipe is an integrated and automatic pipeline that can provide users a friendly solution to perform in-depth mining of MeRIP-seq data. It integrates many functional analysis modules, range from basic processing to downstream analysis. All the processes are embedded in Nextflow with Docker support, which ensures high reproducibility and scalability of the analysis. MeRIPseqPipe is particularly suitable for analyzing a large number of samples at once with a simple command. The final output directory is structured based on each step and tool. And visualization reports containing various tables and plots are provided as HTML files. AVAILABILITY AND IMPLEMENTATION: MeRIPseqPipe is freely available at https://github.com/canceromics/MeRIPseqPipe. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Xiaoqiong Bao, Kaiyu Zhu, Xuefei Liu, Zhihang Chen 0002, Ziwei Luo 0001, Qi Zhao 0009, Jian Ren 0002, Zhixiang Zuo |
Bioinform. | 3 |
| 2021 | End-to-End Spelling Correction Conditioned on Acoustic Feature for Code-Switching Speech Recognition
Shuai Zhang 0014, Jiangyan Yi, Zhengkun Tian, Ye Bai 0001, Jianhua Tao 0001, Xuefei Liu, Zhengqi Wen |
Interspeech | 6 |
| 2021 | TDCA-Net: Time-Domain Channel Attention Network for Depression Detection
Cong Cai, Mingyue Niu, Bin Liu 0041, Jianhua Tao 0001, Xuefei Liu |
Interspeech | 5 |
| 2021 | MPN: Multi-scale Progressive Restoration Network for Unsupervised Defect Detection
Xuefei Liu, Kaitao Song, Jianfeng Lu 0003 |
PRCV (2) | 1 |
| 2021 | M6A2Target: a comprehensive database for targets of m6A writers, erasers and readersabstractN6-methyladenosine (m6A) is the most abundant posttranscriptional modification in mammalian mRNA molecules and has a crucial function in the regulation of many fundamental biological processes. The m6A modification is a dynamic and reversible process regulated by a series of writers, erasers and readers (WERs). Different WERs might have different functions, and even the same WER might function differently in different conditions, which are mostly due to different downstream genes being targeted by the WERs. Therefore, identification of the targets of WERs is particularly important for elucidating this dynamic modification. However, there is still no public repository to host the known targets of WERs. Therefore, we developed the m6A WER target gene database (m6A2Target) to provide a comprehensive resource of the targets of m6A WERs. M6A2Target provides a user-friendly interface to present WER targets in two different modules: 'Validated Targets', referred to as WER targets identified from low-throughput studies, and 'Potential Targets', including WER targets analyzed from high-throughput studies. Compared to other existing m6A-associated databases, m6A2Target is the first specific resource for m6A WER target genes. M6A2Target is freely accessible at http://m6a2target.canceromics.org. Shuang Deng, Hongwan Zhang, Kaiyu Zhu, Xingyang Li, Xuefei Liu, Dongxin Lin, Zhixiang Zuo |
Briefings Bioinform. | 7 |
| 2020 | Epileptic Seizure Prediction Based on Region Correlation of EEG SignalabstractThe existing methods of epileptic seizure prediction usually analyze the electroencephalogram (EEG) signals in the time domain, frequency domain or time-frequency domain. Although some good results have been achieved, the research and utilization of spatial information is still insufficient. Moreover, some studies extracted different features for different patients and achieved good results, but these methods are not universal and robust. Different from the previous methods, this paper propose a new feature processing method of EEG signal. All electrode signals on the scalp are considered as a whole, and fusing data from different regions to obtain spatial information. Then the correlation of first derivatives is used to obtain fluctuation information of signal caused by epilepsy, which further enlarge difference of signal in different seizures stages. In addition, we also design a post-processing strategy, which uses time-series information to rectify prediction results, so that the final result is more accurate. Finally, experimental results from the CHBMIT dataset show effectiveness of proposed method and strategy, while the extensive result confirms that our method is superior to several state-of-the-art methods in recent years. Xuefei Liu, Minglei Shu |
CBMS | 1 |
| 2020 | Non-Autoregressive End-to-End TTS with Coarse-to-Fine Decoding
Tao Wang 0074, Xuefei Liu, Jianhua Tao 0001, Jiangyan Yi, Ruibo Fu, Zhengqi Wen |
INTERSPEECH | 2 |
| 2020 | End-to-End Post-Filter for Speech Separation With Deep Attention Fusion FeaturesabstractIn this article, we propose an end-to-end post-filter method with deep attention fusion features for monaural speaker-independent speech separation. At first, a time-frequency domain speech separation method is applied as the pre-separation stage. The aim of pre-separation stage is to separate the mixture preliminarily. Although this stage can separate the mixture, it still contains the residual interference. In order to enhance the pre-separated speech and improve the separation performance further, the end-to-end post-filter (E2EPF) with deep attention fusion features is proposed. The E2EPF can make full use of the prior knowledge of the pre-separated speech, which contributes to speech separation. It is a fully convolutional speech separation network and uses the waveform as the input features. Firstly, the 1-D convolutional layer is utilized to extract the deep representation features for the mixture and pre-separated signals in the time domain. Secondly, to pay more attention to the outputs of the pre-separation stage, an attention module is applied to acquire deep attention fusion features, which are extracted by computing the similarity between the mixture and the pre-separated speech. These deep attention fusion features are conducive to reduce the interference and enhance the pre-separated speech. Finally, these features are sent to the post-filter to estimate each target signals. Experimental results on the WSJ0-2mix dataset show that the proposed method outperforms the state-of-the-art speech separation method. Compared with the pre-separation method, our proposed method can acquire 64.1%, 60.2%, 25.6% and 7.5% relative improvements in scale-invariant source-to-noise ratio (SI-SNR), the signal-to-distortion ratio (SDR), the perceptual evaluation of speech quality (PESQ) and the short-time objective intelligibility (STOI) measures, respectively. Cunhang Fan, Jianhua Tao 0001, Bin Liu 0041, Jiangyan Yi, Zhengqi Wen, Xuefei Liu |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2018 | Application of Temperature Prediction Based on Neural Network in Intrusion Detection of IoTabstractThe security of network information in the Internet of Things faces enormous challenges. The traditional security defense mechanism is passive and certain loopholes. Intrusion detection can carry out network security monitoring and take corresponding measures actively. The neural network-based intrusion detection technology has specific adaptive capabilities, which can adapt to complex network environments and provide high intrusion detection rate. For the sake of solving the problem that the farmland Internet of Things is very vulnerable to invasion, we use a neural network to construct the farmland Internet of Things intrusion detection system to detect anomalous intrusion. In this study, the temperature of the IoT acquisition system is taken as the research object. It has divided which into different time granularities for feature analysis. We provide the detection standard for the data training detection module by comparing the traditional ARIMA and neural network methods. Its results show that the information on the temperature series is abundant. In addition, the neural network can predict the temperature sequence of varying time granularities better and ensure a small prediction error. It provides the testing standard for the construction of an intrusion detection system of the Internet of Things. Xuefei Liu, Chao Zhang 0015, Pingzeng Liu, Maoling Yan, Baojia Wang, Jianyong Zhang, Russell Higgs |
Secur. Commun. Networks | 1 |
| 2017 | A free shape 3d modeling system for creative design based on modified catmull-clark subdivision
Guanghua Tan, Xianyi Zhu, Xuefei Liu |
Multim. Tools Appl. | 3 |