VLDB 2026 Research / reviewers in the wild / expert
Xingfeng Li 0001
dblp:63/9121-1
· DBLP profile ↗
17ranked-venue papers
7as first author
13since 2021 · last 2026
0000-0002-8958-0341ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 6 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PFDBooster: A unified post-image fusion dual-domain boosting paradigm
Shenzhi Li, Xingfeng Li 0001 |
Knowl. Based Syst. | 2 |
| 2025 | Valence-Arousal Emotion Recognition Using a Deep Three-Layer Model with Aural Perceptual Representations
Xingfeng Li 0001 |
IEEE Big Data | 1 |
| 2025 | Phase-Aware Spectrogram Fusion with Dual-Stream Residual Networks for Underwater Acoustic Recognition
Xingfeng Li 0001, Feifei Yu |
IEEE Big Data | 1 |
| 2025 | Who, When, and What: Leveraging the "Three Ws" Concept for Emotion Recognition in Conversation
Xingfeng Li 0001, Tomoki Toda |
INTERSPEECH | 2 |
| 2025 | Speaker-Aware Multi-Task Learning for Speech Emotion Recognition
Xingfeng Li 0001, Tomoki Toda |
INTERSPEECH | 2 |
| 2025 | Advancing Emotion Recognition via Ensemble Learning: Integrating Speech, Context, and Text Representations
Jinyi Mi, Xingfeng Li 0001, Tomoki Toda |
INTERSPEECH | 3 |
| 2025 | Enhanced Speech Emotion Recognition in Noisy Environments: Adaptive Emotion Denoising Diffusion Approach With Iterative Confidence Learning StrategyabstractSpeech emotion recognition (SER) in noisy environments is challenging due to the overlap of emotional cues with background noise. This paper proposes a novel approach to transfer emotional information from clean to noisy speech, ensuring robust recognition even in adverse conditions. First, 3D Multi-Resolution Modulated Filtered Cochleogram features are extracted to capture dynamic emotional information, while delta and delta-delta features enhance emotional dynamics in both time and frequency domains. Additionally, the Emotional BiMamba Encoder, utilizing bidirectional parallel processing through the Multi-time View Bidirectional State Space Model is designed to capture complex emotional patterns across temporal scales, preserving short-and long-term dependencies. Next, the Adaptive Emotion Denoising Diffusion (AEDD) based on a diffusiondenoising probabilistic model, applies confidence filtering to select representative emotional segments and transfer emotional information from clean to noisy speech, addressing noisy emotional data scarcity. Finally, the Iterative Confidence Learning Strategy (ICLS) enhances the classification network (CN) through a twostage learning process, progressively adapting CN to noisy feature distributions while ensuring consistency during diffusion. The experimental results show significant improvements over the state-of-the-art methods on three datasets: on IEMOCAP, WA increases by 5.07%, UA by 5.23%, and WF1 by 5.44%; on CASIA, WA and UA rise by 2.72%, and WF1 by 2.23%; and on EMODB, WA improves by 3.46%, UA by 3.23%, and WF1 by 3.45%. These consistent gains across different signal-to-noise ratios further validate the effectiveness of the proposed method. Yang Liu 0262, Yarong Li, Xiaoqi Yang 0006, Xiaolei Meng, Xingfeng Li 0001, Zhen Zhao 0006 |
IEEE Internet Things J. | 9 |
| 2024 | RNASite: A one-stop tool website that integrates multiple RNA modification site databases and serversabstractRNASite is a comprehensive platform that integrates multiple RNA modification site databases and servers, focusing on RNA modification sites. These modification sites significantly affect RNA's structure, stability, and function, playing a crucial role in epigenetics and gene expression regulation. The website offers both datasets and online RNA modification site identification tools to enhance the recognition and understanding of these modification sites. With advancements in high-throughput sequencing technology and machine learning, RNASite leverages these innovations to provide efficient and cost-effective RNA modification site identification methods. The platform features 17 high-quality datasets covering 11 common RNA modification sites and includes 7 online identification tools. Notably, some of these tools exceed the accuracy of existing models. RNASite is designed to be a convenient and efficient resource for researchers in biochemistry and bioinformatics, facilitating progress in the study of RNA modification sites. The platform is available for free at http://www.bioai-lab.com/RNASite. Xingfeng Li 0001, Qingchen Zhang 0001, Quan Zou 0001, Feifei Cui |
BIBM | 2 |
| 2024 | MF-AED-AEC: Speech Emotion Recognition by Leveraging Multimodal Fusion, Asr Error Detection, and Asr Error CorrectionabstractThe prevalent approach in speech emotion recognition (SER) involves integrating both audio and textual information to comprehensively identify the speaker’s emotion, with the text generally obtained through automatic speech recognition (ASR). An essential issue of this approach is that ASR errors from the text modality can worsen the performance of SER. Previous studies have proposed using an auxiliary ASR error detection task to adaptively assign weights of each word in ASR hypotheses. However, this approach has limited improvement potential because it does not address the coherence of semantic information in the text. Additionally, the inherent heterogeneity of different modalities leads to distribution gaps between their representations, making their fusion challenging. Therefore, in this paper, we incorporate two auxiliary tasks, ASR error detection (AED) and ASR error correction (AEC), to enhance the semantic coherence of ASR text, and further introduce a novel multi-modal fusion (MF) method to learn shared representations across modalities. We refer to our method as MF-AED-AEC. Experimental results indicate that MF-AED-AEC significantly outperforms the baseline model by a margin of 4.1%. Xingfeng Li 0001, Tomoki Toda |
ICASSP | 3 |
| 2024 | Multimodal Fusion of Music Theory-Inspired and Self-Supervised Representations for Improved Emotion Recognition
Xingfeng Li 0001, Tomoki Toda |
INTERSPEECH | 2 |
| 2024 | Hyb_SEnc: An Antituberculosis Peptide Predictor Based on a Hybrid Feature Vector and Stacked Ensemble LearningabstractTuberculosis has plagued mankind since ancient times, and the struggle between humans and tuberculosis continues. Mycobacterium tuberculosis is the leading cause of tuberculosis, infecting nearly one-third of the world's population. The rise of peptide drugs has created a new direction in the treatment of tuberculosis. Therefore, for the treatment of tuberculosis, the prediction of anti-tuberculosis peptides is crucial. This paper proposes an anti-tuberculosis peptide prediction method based on hybrid features and stacked ensemble learning. First, a random forest (RF) and extremely randomized tree (ERT) are selected as first-level learning of stacked ensembles. Then, the five best-performing feature encoding methods are selected to obtain the hybrid feature vector, and then the decision tree and recursive feature elimination (DT-RFE) are used to refine the hybrid feature vector. After selection, the optimal feature subset is used as the input of the stacked ensemble model. At the same time, logistic regression (LR) is used as a stacked ensemble secondary learner to build the final stacked ensemble model Hyb_SEnc. The prediction accuracy of Hyb_SEnc achieved 94.68% and 95.74% on the independent test sets of AntiTb_MD and AntiTb_RD, respectively. Xiuhao Fu, Xiaofeng Zang, Xingfeng Li 0001, Qingchen Zhang 0001, Quan Zou 0001, Feifei Cui |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2023 | Emotion Awareness in Multi-utterance Turn for Improving Emotion Prediction in Multi-Speaker Conversation
Xingfeng Li 0001, Tomoki Toda |
INTERSPEECH | 2 |
| 2023 | Music Theory-Inspired Acoustic Representation for Speech Emotion RecognitionabstractThis research presents a music theory-inspired acoustic representation (hereafter, MTAR) to address improved speech emotion recognition. The recognition of emotion in speech and music is developed in parallel, yet a relatively limited understanding of MTAR for interpreting speech emotions is involved. In the present study, we use music theory to study representative acoustics associated with emotion in speech from vocal emotion expressions and auditory emotion perception domains. In experiments assessing the role and effectiveness of the proposed representation in classifying discrete emotion categories and predicting continuous emotion dimensions, it shows promising performance compared with extensively used features for emotion recognition based on the spectrogram, Mel-spectrogram, Mel-frequency cepstral coefficients, VGGish, and the large baseline feature sets of the INTERSPEECH challenges. This proposal opens up a novel research avenue in developing a computational acoustic representation of speech emotion via music theory. Xingfeng Li 0001, Desheng Hu, Qingchen Zhang 0001, Zhengxia Wang, Masashi Unoki, Masato Akagi |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2019 | The Contribution of Acoustic Features Analysis to Model Emotion Perceptual Process for Language DiversityabstractThe multi-layered perceptual process of emotion in human speech plays an essential role in the field of affective computing for underlying a speaker’s state. However, a comprehensive process analysis of emotion perception is still challenging due to the lack of powerful acoustic features allowing accurate inference of emotion across speaker and language diversities. Most previous research works study acoustic features mostly using Fourier transform, short time Fourier transform or linear predictive coding. Even though these features may be useful for stationary signal within short frames, they may not capture the localized event adequately as speech transmits emotion information dynamically over time. This case introduces a set of acoustic features via wavelet transform analysis of the speech signal, and specifically, models the perceptual process of emotion for language diversity. For this aim, the proposed features are analyzed in a three-layer emotion perception model across multiple languages. Experiments show that the proposed acoustic features significantly enhance the perceptual process of emotion and render a better result in multilingual emotion recognition when compared it to the widely used prosodic and spectral features, as well as their combination in literature. Xingfeng Li 0001, Masato Akagi |
INTERSPEECH | 1 |
| 2019 | Improving multilingual speech emotion recognition by combining acoustic features in a three-layer model
Xingfeng Li 0001, Masato Akagi |
Speech Commun. | 1 |
| 2018 | A Three-Layer Emotion Perception Model for Valence and Arousal-Based Detection from Multilingual SpeechabstractAutomated emotion detection from speech has recently shifted from monolingual to multilingual tasks for human-like interaction in real-life where a system can handle more than a single input language. However, most work on monolingual emotion detection is difficult to generalize in multiple languages, because the optimal feature sets of the work differ from one language to another. Our study proposes a framework to design, implement and validate an emotion detection system using multiple corpora. A continuous dimensional space of valence and arousal is first used to describe the emotions. A three-layer model incorporated with fuzzy inference systems is then used to estimate two dimensions. Speech features derived from prosodic, spectral and glottal waveform are examined and selected to capture emotional cues. The results of this new system outperformed the existing state-of-the-art system by yielding a smaller mean absolute error and higher correlation between estimates and human evaluators. Moreover, results for speaker independent validation are comparable to human evaluators. Xingfeng Li 0001, Masato Akagi |
INTERSPEECH | 1 |
| 2016 | Multilingual Speech Emotion Recognition System Based on a Three-Layer Model
Xingfeng Li 0001, Masato Akagi |
INTERSPEECH | 1 |