Dongdong Li 0003

dblp:14/5457-3 · DBLP profile ↗
← Back
77ranked-venue papers
22as first author
51since 2021 · last 2027
0000-0002-1880-8054ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 50 · 14 first-author · 30 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 1 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 6 since 2021Databases, data management, data science and information retrieval · 6 · 3 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 first-authorComputer networks · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2027 Dense multi-scale graph token transformer for local-global graph representation learning
Dongdong Li 0003, Jinchen Han, Zhe Wang 0002
Expert Syst. Appl.1
2026 NEURAL-VOX: NEURal auditory language decoding for voice and text reconstruction
Zhishuo Jin, Dongdong Li 0003, Qin Zhou 0002, Zhe Wang 0002
Neural Networks2
2026 TCFnet: Temporal-frequency hybrid MetaFormer with multi-stage training for neural tracking of speech
Dongdong Li 0003, Shengyao Huang, Zhongliang Zeng, Zhishuo Jin
Pattern Recognit.1
2026 PDAug: Population-based Dynamic Audio Augmentation for imbalanced data depression detection
Dongdong Li 0003, Zhouhang Wang, Zhe Wang 0002
Speech Commun.1
2026 Fine-Grained Emotion Adaptive Alignment Network for Image Emotion Distribution Transfer
Jing Zhang 0041, Jixiang Zhu, Yumo Kang, Dongdong Li 0003, Zhe Wang 0002
IEEE Trans. Affect. Comput.4
2026 FAD3QN: A Brain-Inspired Deep Reinforcement Learning Model for Speech Depression Detection
abstract
In recent years, the high prevalence and severity of depression have highlighted the urgent need for early detection. Depressed patients exhibit noticeable emotional changes in their speech, but existing detection methods face significant challenges in modeling emotion perception mechanisms. In this article, inspired by the knowledge of reinforcement learning neuroscience and the theory of emotion perception in the limbic system, we propose a brain-inspired model frontal-amygdala double dueling deep Q network for depression detection based on speech. The model simulates the frontal lobe and amygdala-centred brain mechanisms for emotion perception through the reinforcement learning framework of double dueling deep Q-networks, and embeds the skip-connected 1-D convolutional neural network and bidirectional long short-term memory network neural networks to simulate the emotion perception process in the limbic system. In addition, we designed an adaptively tuned reward function to address the data imbalance in depression detection, and incorporated an additional step-size penalty factor to limit the number of incorrect decisions made by the agent during the early stages of training. Experimental results across multiple datasets demonstrate the effectiveness and generalizability of our approach. Meanwhile, the relevant ablation experiments conducted in this article validate the key roles of the frontal and limbic systems in the reward learning and emotion perception process, as well as the importance of the adaptive reward function in solving the data imbalance problem.
Dongdong Li 0003, Zhe Wang 0002, Yichao Yin
IEEE Trans. Comput. Soc. Syst.1
2025 Ipvar: Advancing Pathogenicity Prediction Via Hierarchical Fusion of Structure Foundation Models Alphafold 3 and Esm C
abstract
Interpreting the functional consequences of coding variants remains a central challenge in human genetics, particularly given the clinical importance of distinguishing pathogenic mutations from benign variation. Here we present IPVAR, a deep learning framework that uniquely integrates tertiary protein structural features predicted by both ESM C and AlphaFold 3, alongside established conservation metrics, to advance variant pathogenicity prediction. IPVAR leverages a hierarchical crossattention mechanism to capture both global and fine-grained structural dependencies between complementary representations, and incorporates an adaptive modality weighting strategy to dynamically balance information from each protein structure model. Comprehensive benchmarking demonstrates that IPVAR substantially outperforms state-of-the-art methods, including those based solely on sequence annotations or individual structural predictors, achieving an area under the ROC curve (AUC) of 0.9838 on the ClinVar dataset and 0.9202 on an independent Mendelian disease variant cohort. Ablation studies further confirm that both the multi-model integration and advanced fusion modules are critical to the model's superior performance. These findings establish IPVAR as a new benchmark for the functional interpretation of genomic variants, and highlight the value of integrating diverse structural foundation models to improve clinical variant assessment.
Hanwen Huang, Ziquan Bao, Yingzhuo Wang, Qin Zhou 0002, Ting Xiao 0002, Qian Zhang 0068, Dongdong Li 0003, Hai Yang 0002
BIBM9
2025 Subtype-Former: A Deep Learning Approach for Cancer Subtype Discovery with Multi-Omics Data
abstract
Cancer is heterogeneous, affecting the precise approach to personalized treatment. Accurate subtyping can lead to better survival rates for cancer patients. High-throughput technologies provide multiple omics data for cancer subtyping. This study proposed Subtype-Former, a deep learning method based on MLP and Transformer Block, to extract the lowdimensional representation of the multi-omics data. K-means and Consensus Clustering are also used to achieve accurate subtyping results. We compared Subtype-Former with the other state-of-the-art subtyping methods across the TCGA 10 cancer types. We found that Subtype-Former can perform better on the benchmark datasets of more than 5000 tumors based on the survival analysis. In addition, Subtype-Former also achieved outstanding results in pan-cancer subtyping, which can help analyze the commonalities and differences across various cancer types at the molecular level. Finally, we applied Subtype-Former to the TCGA 10 types of cancers. We identified 50 essential biomarkers, which can be used to study targeted cancer drugs and promote the development of cancer treatments in the era of precision medicine.
Hanwen Huang, Yuhang Sheng, Dongdong Li 0003, Jing Zhang 0041, Hai Yang 0002
BIBM4
2025 TFCAF-Net: A Time-Frequency Co-Attentive Fusion Network for Speech-Based Depression Detection
abstract
Speech contains important cues related to mental health status, making it a valuable signal source for automatic depression detection. However, existing approaches often exhibit insufficient modeling of the coupling between temporal and spectral dynamics in speech, limiting the ability to capture their complementary information associated with depressive symptoms. This paper presents a Time-Frequency Co-Attentive Fusion Network (TFCAF-Net) that jointly models temporal and spectral representations of speech to enhance detection reliability. The proposed architecture consists of two attention-enhanced branches: a temporal branch that captures sequential dynamics in speech and a spectral branch that focuses on frequency-domain characteristics. A Time-Frequency Co-Attentive Pooling (TFCAP) module is further introduced to adaptively integrate the two feature streams through attention, enabling mutual guidance during feature fusion. Experiments conducted on the CMDC and DAIC_WOZ datasets show that TFCAF-Net achieves competitive performance compared with existing baselines, validating the effectiveness of the proposed attention-based fusion strategy for capturing complementary time-frequency information relevant to depression detection.
Dongdong Li 0003
BIBM2
2025 Depression Detection Based on Self-Supervised Pretrained Speech Models
abstract
Depression is a global mental-health challenge, demanding objective and scalable screening tools. Speech—noninvasive, universally available, and rich in affective cues—has emerged as a promising biomarker. While early systems relied on hand-crafted features, the advent of self-supervised representation learning now enables high-capacity models to capture subtle acoustic signatures of mood disturbance across languages and cultures. Thus, this study proposes a modular, datasetagnostic framework that investigates the impact of three key design choices: the variation of self-supervised speech encoder, the feature extraction strategy, and the family of downstream classifier. By rigorously evaluating these components within a unified experimental protocol, we illuminate the factors that govern generalizable depression detection, highlighting performance differences across various model and strategies combinations.
Peite Guo, Dongdong Li 0003
BIBM4
2025 Multimodal Fusion for EEG Emotion Recognition in Music with a Multi-Task Learning Framework
abstract
This paper proposes a novel EEG-based emotion recognition approach for music, employing a two-stage training framework that integrates emotion representations from music, lyrics, and EEG. First, a modality-specific feature extraction strategy fine-tunes encoders for music and lyrics to extract emotion-related features, while the EEG encoder is fine-tuned to capture identity-related features. The second stage applies a local-to-global fusion strategy, merging EEG and music features for detailed, time-aligned modeling, and integrating global semantic representations from lyrics. Additionally, a multi-task learning framework is utilized to disentangle emotion-related and identity-related information in EEG, enabling robust, subject-independent emotion recognition. Experimental results on the EREMUS dataset show a significant improvement over baseline models, achieving a total score of 51.97%, demonstrating the effectiveness of multimodal integration and our proposed framework for emotion recognition in music.
Shengyao Huang, Zhishuo Jin, Dongdong Li 0003, Jinchen Han, Xie Tao 0004
ICASSP3
2025 TFCTL: Time-Frequency Calibrated Transfer Learning for cross domain depression detection
Dongdong Li 0003, Zuo Yang, Zhe Wang 0002
Eng. Appl. Artif. Intell.1
2025 Resource-efficient cross-subject emotion recognition from electroencephalogram via spiking domain discriminators
Dongdong Li 0003, Shengyao Huang, Yujun Shen, Zhe Wang 0002
Eng. Appl. Artif. Intell.1
2025 Trans-Driver: A Deep Learning Approach for Cancer Driver Gene Discovery With Multi-Omics Data
abstract
Driver genes play a crucial role in the growth of cancer cells. Accurate identification of cancer driver genes is essential for deepening our understanding of cancer pathogenesis and facilitating the development of cancer therapies and drug-targeted driver genes. However, the diversity and complexity of multi-omics data still make cancer driver identification highly challenging. In this study, we propose Transformer-Driver (Trans-Driver), a deep supervised learning method based on a novel transformer architecture, which integrates multi-omics data to learn the differences and associations between different omics modalities for cancer driver discovery. Trans-Driver introduces a kernel-based multi-head self-attention mechanism with gated residual connections, as well as a Dynamic Tanh (DyT) normalization function, to enhance the integration and modeling of heterogeneous multi-omics features. Compared with other state-of-the-art driver gene identification methods, Trans-Driver achieved excellent performance on TCGA, CGC, and PCAWG datasets. Among approximately 20,000 protein-coding genes, Trans-Driver reported 269 candidate driver genes, of which 132 genes (about 49.1%) were included in the gold standard CGC dataset. Feature contribution analysis further demonstrated that integrating multi-omics data improved performance compared to using somatic mutation data alone. Finally, detailed analysis revealed that the candidate drivers are clinically meaningful, demonstrating the practical value of Trans-Driver.
Hai Yang 0002, Zhenbei Yang, Lei Zhang 0224, Yijing Yang, Dongdong Li 0003, Jing Zhang 0041, Zhe Wang 0002
IEEE Trans. Comput. Biol. Bioinform.6
2025 Double Confidence Calibration Focused Distillation for Task-Incremental Learning
abstract
Task-incremental learning methods that adopt knowledge distillation face two significant challenges: confidence bias and knowledge loss. These challenges make it difficult to effectively balance the stability and plasticity of the network in the incremental learning process. In this article, we propose double confidence calibration focused distillation (DCCFD) to address these challenges. We introduce intratask and intertask confidence calibration (ECC) modules that can mitigate network overconfidence during incremental learning and reduce the degree of feature representation bias. We also propose a focused distillation (FD) module that can alleviate the problem of knowledge loss during the task increment process, improving model stability without reducing plasticity. Experimental results on the CIFAR-100, TinyImageNet, and CORE-50 datasets demonstrate the effectiveness of our method, with performance that matches or exceeds the state of the art. Furthermore, our method can be used as a plug-and-play module to consistently improve class-incremental learning methods.
Zhiling Fu, Zhe Wang 0002, Chengwei Yu, Xinlei Xu, Dongdong Li 0003
IEEE Trans. Neural Networks Learn. Syst.5
2025 Neuron Perception Inspired EEG Emotion Recognition With Parallel Contrastive Learning
abstract
Considerable interindividual variability exists in electroencephalogram (EEG) signals, resulting in challenges for subject-independent emotion recognition tasks. Current research in cross-subject EEG emotion recognition has been insufficient in uncovering the shared neural underpinnings of affective processing in the human brain. To address this issue, we propose the parallel contrastive multisource domain adaptation (PCMDA) model, inspired by the neural representation mechanism in the ventral visual cortex. Our model employs a neuron-perception-inspired contrastive learning architecture for EEG-based emotion recognition in subject-independent scenarios. A two-stage alignment methodology is employed for the purpose of aligning numerous source domains with the target domain. This approach integrates a parallel contrastive loss (PCL) which simulates the self-supervised learning mechanism inherent in the neural representation of the human brain. Furthermore, a self-attention mechanism is integrated to extract emotion weights for each frequency band. Extensive experiments were conducted on three publicly available EEG emotion datasets, SJTU emotion EEG dataset (SEED), database for emotion analysis using physiological signals (DEAP), and finer-grained affective computing EEG dataset (FACED), to evaluate our proposed method. The results demonstrate that the PCMDA effectively utilizes the unique EEG features and frequency band information of each subject, leading to improved generalization across different subjects in comparison to other methods.
Dongdong Li 0003, Shengyao Huang, Zhe Wang 0002
IEEE Trans. Neural Networks Learn. Syst.1
2025 Local-Global Geometric Information and View Complementarity Introduced Multiview Metric Learning
abstract
Geometry studies the spatial structure and location information of objects, providing a priori knowledge and intuitive explanation for classification methods. Considering samples from a geometric perspective offers a novel approach to understanding their information. In this article, we propose a method called local-global geometric information and view complementarity introduced multiview metric learning (GIVCMML). Our method effectively exploits the geometric information of multiview samples. The learned metric space retains the geometric relations of samples and makes them more separable. First, we propose the global geometrical constraint in the maximum margin criterion framework. By maximizing the distance between class centers in the metric space, we ensure that samples from different classes are well separated. Second, to maintain the manifold structure of the original space, we build an adjacency matrix that contains the sample label information. This helps explore the local geometric information of sample pairs. Finally, to better mine the complementary information of multiview samples, GIVCMML maximizes the correlation between each view in the metric space. This enables each view to adaptively learn from the others and explore the complementary information between views. We extensively evaluate the effectiveness of our method on real-world datasets. The experimental results demonstrate that GIVCMML achieves competitive performance compared with multiview metric learning (MvML) methods.
Xinlei Xu, Zhe Wang 0002, Shuangyan Ren, Saisai Niu, Dongdong Li 0003
IEEE Trans. Neural Networks Learn. Syst.5
2024 Optimizing Clinical Depression Detection: Extracting Depression-Specific Feature Sets Using spFSR to Enhance Speech-Based Diagnosis
abstract
Accurate detection of depression through speech analysis offers a promising non-invasive approach for early diagnosis and intervention. However, the high dimensionality and complexity of speech features present significant challenges in identifying the most relevant features for depression detection. This study applies the spFSR (Feature Selection and Ranking via Simultaneous Perturbation Stochastic Approximation) technique to a comprehensive 2,268-dimensional speech feature set, focusing on selecting features specifically relevant to depression. The effectiveness of the spFSR method is evaluated using two well-known datasets: DAIC-WOZ and CMDC. The selected feature set was assessed across various machine learning models, demonstrating substantial improvements in key performance metrics on both datasets. The results indicate that the spFSR method effectively optimizes feature selection for depression detection, leading to more robust and accurate predictive models across different datasets. Our study find that the top 10 features for effective speech-based depression detection include spectral features (e.g., spectral flux, entropy, flatness), fundamental frequency metrics (e.g., lowest percentile), periodic features (e.g., jitter), and MFCC attributes (e.g., segment length, skewness).
Wenhui Guo, Binxiao Chen, Manyue Gu, Dongdong Li 0003, Hai Yang 0002
BIBM5
2024 Emotion embedding framework with emotional self-attention mechanism for speaker recognition
Dongdong Li 0003, Jinlin Liu, Hai Yang 0002, Zhe Wang 0002
Expert Syst. Appl.1
2024 DRSCDM: A Novel Density-Related Clustering for Complex High-Dimensional Data Streams
abstract
The proliferation of high-dimensional complex data in various fields such as multimedia, social media, and sensor networks has led to an increasing demand for real-time clustering algorithms. This article presents a novel two-stage approach for complex data streams. In the online stage, angular margin are introduced to constrain the mapping of input data, enhancing the directional characteristics of the resulting data representation. In the offline stage, we propose a unique clustering approach grounded in angular density to uncover spatial relationships within the data. This approach utilizes two distinct strategies for angular density clustering. Neighbor Selection based on Angular Relations define the angular density, which significantly enhances the algorithm’s discriminative ability. Density-Priority Cluster Selection strategy determines the generation of clusters, ensuring the reliability of clustering. We also introduce a novel data expiration mechanism that optimizes computational costs and memory usage by discarding data objects from stable clusters. Experimental evaluations on four diverse datasets, including speaker diarization and video face clustering tasks, demonstrate the superior performance of our proposed method over state-of-the-art online clustering techniques. Furthermore, our method achieves comparable performance to offline clustering methods, highlighting its effectiveness and efficiency in real-time clustering applications. The source code for the proposed algorithms is accessible athttps://github.com/sssssuda/DRSCDM.
Dongdong Li 0003, Yihan Fan, Zhe Wang 0002
IEEE Trans. Circuits Syst. Video Technol.1
2024 Brain Emotion Perception Inspired EEG Emotion Recognition With Deep Reinforcement Learning
abstract
Inspired by the well-known Papez circuit theory and neuroscience knowledge of reinforcement learning, a double dueling deep Q network (DQN) is built incorporating the electroencephalogram (EEG) signals of the frontal lobe as prior information, which is named frontal lobe double dueling DQN (FLD3QN). The framework of FLD3QN is constructed in accord with the brain emotion mechanism which takes the frontal lobe and the thalamus as the core, in which the part of the Papez circuit is simulated by the bifrontal lobe residual convolution neural network (BiFRCNN). Moreover, a step penalty factor is designed to constrain the number of mistakes of the agent. The ablation studies results on the public EEG emotion dataset DEAP verified the important roles of the frontal lobe and the Papez circuit in modeling the procedure of learning rewards during the perception of emotions, with a great increase in the average accuracies by 25.24% and 23.31% in valence and arousal dimensions.
Dongdong Li 0003, Zhe Wang 0002, Hai Yang 0002
IEEE Trans. Neural Networks Learn. Syst.1
2023 Mixed Entropy Down-Sampling based Ensemble Learning for Speech Emotion Recognition
abstract
The strength of emotion at different positions in a speech is strong or weak, and the weak parts with unclear emotions will bring noise to the model. We propose a boosting ensemble learning method based on mixed entropy down-sampling to effectively select emotionally salient segments to improve the classifier's performance. An independent Convolutional Neural Network (CNN) model is trained in each iteration of ensemble learning. These CNN models form an ensemble classifier, which improves the generalization ability by synthesizing all the learned results of down-sampling, making emotion recognition more accurate. We also introduce the concept of Mixed Information Entropy (MIE), which consists of Emotional Certainty Entropy (ECE) and Structural Distribution Entropy (SDE). ECE measures the emotional confusion of segments, while SDE measures the stability of segments in deep feature space. During the iteration, the deep features are obtained from the last fully connected layer of the model and down-sampled according to the weighted sum of confidence and MIE. The selected segments with stronger emotions are used for the next iteration. Our method is 3.77% higher on WA and 2.37% higher on UA than the naive CNN model on the IEMOCAP dataset.
Zhengji Xuan, Dongdong Li 0003, Zhe Wang 0002, Hai Yang 0002
IJCNN2
2023 Multi-level Feature Joint Learning Methods for Emotional Speaker Recognition
abstract
In the real scene, changes in speaker features caused by different emotional states have a great impact on the performance of speaker recognition. To improve the robustness of the speaker recognition system, the existing emotional speaker recognition technologies tend to cascade different models, ignoring the frame-level acoustic features and the segment-level discourse habits feature. To this end, we combine frame- and segment-level features in different ways to build a robust recognition system for emotional speakers. The frame-level features and segment-level features are jointly learned to retain emotional information and speaker information. Four joint learning methods, namely, Joint in series, Joint in Parallel, Joint under the guidance, and Joint with Original Feature, are discussed to explore the correlations between fragment-level features and frame-level features. The experimental results illustrate that the speaker feature will change greatly in different emotional states. Compared with the accuracy of 90.95% by x-vector, the proposed methods of Joint in parallel and Joint with Original Features can achieve the accuracy of 95.06% and 94.67% respectively for emotional speaker recognition in the experiment on Mandarin Affective Speech Corpus (MASC). Our findings provide a novel aspect to improve speaker recognition robustness.
Zhongliang Zeng, Dongdong Li 0003, Zhe Wang 0002, Hai Yang 0002
IJCNN2
2023 From multi-omics data to the cancer druggable gene discovery: a novel machine learning-based approach
abstract
The development of targeted drugs allows precision medicine in cancer treatment and optimal targeted therapies. Accurate identification of cancer druggable genes helps strengthen the understanding of targeted cancer therapy and promotes precise cancer treatment. However, rare cancer-druggable genes have been found due to the multi-omics data's diversity and complexity. This study proposes deep forest for cancer druggable genes discovery (DF-CAGE), a novel machine learning-based method for cancer-druggable gene discovery. DF-CAGE integrated the somatic mutations, copy number variants, DNA methylation and RNA-Seq data across ˜10 000 TCGA profiles to identify the landscape of the cancer-druggable genes. We found that DF-CAGE discovers the commonalities of currently known cancer-druggable genes from the perspective of multi-omics data and achieved excellent performance on OncoKB, Target and Drugbank data sets. Among the ˜20 000 protein-coding genes, DF-CAGE pinpointed 465 potential cancer-druggable genes. We found that the candidate cancer druggable genes (CDG) are clinically meaningful and divided the CDG into known, reliable and potential gene sets. Finally, we analyzed the omics data's contribution to identifying druggable genes. We found that DF-CAGE reports druggable genes mainly based on the copy number variations (CNVs) data, the gene rearrangements and the mutation rates in the population. These findings may enlighten the future study and development of new drugs.
Hai Yang 0002, Lipeng Gan, Rui Chen 0021, Dongdong Li 0003, Jing Zhang 0041, Zhe Wang 0002
Briefings Bioinform.4
2023 InDEP: an interpretable machine learning approach to predict cancer driver genes from multi-omics data
abstract
Cancer driver genes are critical in driving tumor cell growth, and precisely identifying these genes is crucial in advancing our understanding of cancer pathogenesis and developing targeted cancer drugs. Despite the current methods for discovering cancer driver genes that mainly rely on integrating multi-omics data, many existing models are overly complex, and it is difficult to interpret the results accurately. This study aims to address this issue by introducing InDEP, an interpretable machine learning framework based on cascade forests. InDEP is designed with easy-to-interpret features, cascade forests based on decision trees and a KernelSHAP module that enables fine-grained post-hoc interpretation. Integrating multi-omics data, InDEP can identify essential features of classified driver genes at both the gene and cancer-type levels. The framework accurately identifies driver genes, discovers new patterns that make genes as driver genes and refines the cancer driver gene catalog. In comparison with state-of-the-art methods, InDEP proved to be more accurate on the test set and identified reliable candidate driver genes. Mutational features were the primary drivers for InDEP's identifying driver genes, with other omics features also contributing. At the gene level, the framework concluded that substitution-type mutations were the main reason most genes were identified as driver genes. InDEP's ability to identify reliable candidate driver genes opens up new avenues for precision oncology and discovering new biomedical knowledge. This framework can help advance cancer research by providing an interpretable method for identifying cancer driver genes and their contribution to cancer pathogenesis, facilitating the development of targeted cancer drugs.
Hai Yang 0002, Yijing Yang, Dongdong Li 0003, Zhe Wang 0002
Briefings Bioinform.4
2023 DFSGAN: Introducing editable and representative attributes for few-shot image generation
Mengping Yang, Saisai Niu, Zhe Wang 0002, Dongdong Li 0003, Wenli Du
Eng. Appl. Artif. Intell.4
2023 Personalized Federated Continual Learning for Task-Incremental Biometrics
abstract
In the age of Internet of Things where information is explosively growing, people pay more attention on personal privacy. In the real-world task-incremental scenario for biometrics, every edge device faces continuous task flows of private data without communication with others. security and performance are the primary concerns in identity authentication, and federated continual learning (FCL) is a promising solution. In this article, we design a personalized FCL framework to solve the problem of sequential identification in every distributed device. For each client, we create an adaptive continual metalearning model called continual task-distillation-based adaptive model-agnostic metalearning (cTD-$\alpha $MAML), aiming to align the gradients of previous and new tasks and to make the learning-rate (LR) model learnable. For central aggregation, the server gathers the metainitialization from every local update and allocates the updated global metainitialization to clients. We propose an extension of federated average to locally reserve the learnable LR network to realize the personalization of clients. Results prove that in continual learning, our cTD-$\alpha $MAML can learn to learn the seen tasks and avoid catastrophic forgetting. And in FCL, our personalized method realizes the knowledge transferring across clients, meanwhile improving the local performance and reducing the communication cost. In this way, the proposed personalized FCL framework can obtain a biometric template that is able to learn the expression space for new tasks with rapid adaption.
Dongdong Li 0003, Zhe Wang 0002, Hai Yang 0002
IEEE Internet Things J.1
2023 Scalable one-stage multi-view subspace clustering with dictionary learning
Wei Guo 0023, Zhe Wang 0002, Ziqiu Chi, Xinlei Xu, Dongdong Li 0003
Knowl. Based Syst.5
2023 Flexible few-shot class-incremental learning with prototype container
Xinlei Xu, Zhe Wang 0002, Zhiling Fu, Wei Guo 0023, Ziqiu Chi, Dongdong Li 0003
Neural Comput. Appl.6
2023 Knowledge aggregation networks for class incremental learning
Zhiling Fu, Zhe Wang 0002, Xinlei Xu, Dongdong Li 0003, Hai Yang 0002
Pattern Recognit.4
2023 Identity Retention and Emotion Converted StarGAN for low-resource emotional speaker recognition
Dongdong Li 0003, Zhe Wang 0002, Hai Yang 0002
Speech Commun.1
2023 Multiple Kernel Subspace Learning for Clustering and Classification
abstract
In the face of high-dimensional and complex data, effective subspace can preserve specific statistical properties and provide an appropriate representation of data, which generally facilitates the underlying tasks such as clustering or classification. Meanwhile, multiple kernel learning is a technique to combine multiple kernels from different feature spaces effectively. Thus, by incorporating multiple kernels into the process of subspace learning, different feature spaces can be projected into a unified subspace. This paper proposes the Multiple Kernel Subspace Learning (MKSL) for embedding the original space into a unified subspace. Multiple kernels of different feature spaces are combined by MKSL in the process of learning, which can extend the suitability for various applications. Moreover, to generate the optimal combination kernel of subspace learning, we propose a two-step iteration strategy to learn the appropriate kernel weights and transformation matrix of projecting simultaneously. Furthermore, our proposed formulation of MKSL can introduce different prior knowledge such as class information and neighborhood relationships. Thus it is competent to the unsupervised learning, semi-supervised learning, and supervised learning. Extensive experiments are conducted on diverse datasets, and the performances are comprehensively evaluated on different tasks. The experimental results indicate that the proposed algorithm is outstanding in unsupervised clustering task and effective in supervised and semi-supervised classification tasks.
Ziqiu Chi, Zhe Wang 0002, Bolu Wang, Zhongli Fang, Zonghai Zhu, Dongdong Li 0003, Wenli Du
IEEE Trans. Knowl. Data Eng.6
2023 Multilabel Convolutional Network With Feature Denoising and Details Supplement
abstract
In multilabel images, the changeable size, posture, and position of objects in the image will increase the difficulty of classification. Moreover, a large amount of irrelevant information interferes with the recognition of objects. Therefore, how to remove irrelevant information from the image to improve the performance of label recognition is an important problem. In this article, we propose a convolutional network based on feature denoising and details supplement (FDDS) to address this issue. In FDDS, we first design a cascade convolution module (CCM) to collect spatial details of upper features, in order to enhance the information expression of features. Second, the feature denoising module (FDM) is further put forward to reallocate the weight of the feature semantic area, in order to enrich the effective semantic information of the current feature and perform denoising operations on object-irrelevant information. Experimental results show that the proposed FDDS outperforms the existing state-of-the-art models on several benchmark datasets, especially for complex scenes.
Tianhao Gu, Zhe Wang 0002, Zhongli Fang, Zonghai Zhu, Hai Yang 0002, Dongdong Li 0003, Wenli Du
IEEE Trans. Neural Networks Learn. Syst.6
2023 Frame-Level Teacher-Student Learning With Data Privacy for EEG Emotion Recognition
abstract
Recently, electroencephalogram (EEG) emotion recognition has gradually attracted a lot of attention. This brief designs a novel frame-level teacher-student framework with data privacy (FLTSDP) for EEG emotion recognition. The framework first proposes a teacher-student network without prior professional information for automated filtering of useful frame-level features by a gated mechanism and extracting high-level features by using knowledge distillation to capture the results of EEG emotion recognition from a teacher network and student networks. Then, the results from subnetworks are integrated by using the novel decision module, which, motivated by the voting mechanism, adjusts the composition of feature vectors and improves the weight of accurate prediction to optimize the integration effect. During training, an innovative data privacy protection mechanism is applied for avoiding data sharing, where each student network only inherits weights from all trained networks and does not inherit the training dataset. Here, the framework can be repeatedly optimized and improved by only training the next student subnetwork on new EEG signals. Experimental results show that our framework improves the accuracy of EEG emotion recognition by more than 5% and gets state-of-the-art performance for EEG emotion recognition in the subject-independent mode.
Tianhao Gu, Zhe Wang 0002, Xinlei Xu, Dongdong Li 0003, Hai Yang 0002, Wenli Du
IEEE Trans. Neural Networks Learn. Syst.4
2022 CFC: a Cascade Forest approach to discover Cancer driver genes using multi-omics data
abstract
With the development of next-generation sequencing technology, massive genomic data has been generated, primarily encouraging research on cancer driver genes. Many bioinformatics methods were proposed to identify driver genes. However, the results of driver gene identification a mong these methods show considerable differences. It is still challenging to obtain a comprehensive catalog of cancer drivers. Although current methods have greatly promoted the development of driver genes, few methods can integrate the identification results of existing methods. To solve such problems in cancer driver genes research, we proposed a cascade forest model to discover cancer driver genes(CFC) that can integrate multi-omics data and annotation scores from different cancer driver gene identification algorithms. The proposed method got precise results for 33 cancer types and Pan-cancer. The CFC framework identified 275 driver genes in Pan-cancer, of which 179 were included in the Gold standard. The identified genes were enriched i n t he principal cancer signaling pathways.
Lei Zhang 0224, Yijing Yang, Zhe Wang 0002, Dongdong Li 0003, Hai Yang 0002
BIBM4
2022 Deep Spatio-Temporal Mutual Learning for EEG Emotion Recognition
abstract
EEG emotion recognition is an essential area of brain-computer interface(BCI). Because of the low signal-to-noise ratio (SNR) and the uncertainty of the relationship between channels, it is arduous to mine the spatial and temporal information of EEG, especially through a single data representation method. Nowadays, several studies have applied knowledge distillation to the field of emotion recognition. However, traditional knowledge distillation requires a more powerful teacher model, which is time-consuming and needs massive storage space. In order to solve the above problems, in this paper, we propose a novel deep spatio-temporal mutual learning architecture named MLBNet for EEG emotion recognition, which is composed of temporal biased feature learner and spatial biased feature learner. The two components can learn well from chain-like data and matrix-like data respectively, and are trained collaboratively to mimic the predicted probability of each other. By the proposed architecture, we can improve the performance of EEG emotion recognition simply and effectively. To evaluate the validity of proposed method, we performed subject-dependent binary-class and four-class emotion identification tasks on DEAP dataset. The average result of the 10-fold cross-validation is considered as the final result. The MLBNet achieves 98.72% accuracy on valence and 98.85% accuracy on arousal respectively, and 98.32% accuracy on four-class classification tasks. To our best knowledge, our model demonstrates a better performance than the state-of-the-art models with the identical settings.
Wenqing Ye, Haokun Zhang, Zhuolin Zhu, Dongdong Li 0003
IJCNN5
2022 Multi-attention mutual information distributed framework for few-shot learning
Zhe Wang 0002, Pingchuan Ma 0009, Ziqiu Chi, Dongdong Li 0003, Hai Yang 0002, Wenli Du
Expert Syst. Appl.4
2022 Semi-supervised multiple empirical kernel learning with pseudo empirical loss and similarity regularization
abstract
Multiple empirical kernel learning (MEKL) is a scalable and efficient supervised algorithm based on labeled samples. However, there is still a huge amount of unlabeled samples in the real-world application, which are not applicable for the supervised algorithm. To fully utilize the spatial distribution information of the unlabeled samples, this paper proposes a novel semi-supervised multiple empirical kernel learning (SSMEKL). SSMEKL enables multiple empirical kernel learning to achieve better classification performance with a small number of labeled samples and a large number of unlabeled samples. First, SSMEKL uses the collaborative information of multiple kernels to provide a pseudo labels to some unlabeled samples in the optimization process of the model, and SSMEKL designs pseudo-empirical loss to transform learning process of the unlabeled samples into supervised learning. Second, SSMEKL designs the similarity regularization for unlabeled samples to make full use of the spatial information of unlabeled samples. It is required that the output of unlabeled samples should be similar to the neighboring labeled samples to improve the classification performance of the model. The proposed SSMEKL can improve the performance of the classifier by using a small number of labeled samples and numerous unlabeled samples to improve the classification performance of MEKL. In the experiment, the results on four real-world data sets and two multiview data sets validate the effectiveness and superiority of the proposed SSMEKL.
Wei Guo 0023, Zhe Wang 0002, Menghao Ma, Lilong Chen, Hai Yang 0002, Dongdong Li 0003, Wenli Du
Int. J. Intell. Syst.6
2022 Boundary-based Fuzzy-SVDD for one-class classification
abstract
Support Vector Data Description (SVDD) is an extremely hot topic issue in One-Class Classification (OCC), which has displayed outstanding performance in dealing with many novelty detection problems. However, SVDD just takes the data description by the kernel-based distance among each instance into consideration rather than considering the distribution of the data. Therefore, Fuzzy Support Vector Data Description (Fuzzy-SVDD) has been developed to distribute a fuzzy membership to each input sample so that different samples cause different contributions to classification boundary. The majority of the methods in Fuzzy-SVDD are based on the sample density, but there are remaining two problems. These density-based Fuzzy-SVDD methods would decrease the contribution of support vectors (SVs) in low densities. What is more, these methods cannot get a precise density when there are few target samples. These two problems would lead to a poor classification boundary. To overcome these drawbacks, a novel method called Boundary-based Fuzzy-SVDD (BF-SVDD) is proposed in this paper. BF-SVDD uses a new definition called local–global center distance to search for the samples near the boundary. Then, it enhances fuzzy memberships of these samples because they carry more significant information for the decision boundary than other data. The contribution of this paper can be summarized into three main points. First a novel concept called local–global center distances is proposed to find the SVs better. Second, fuzzy memberships with local–global center distance make SVs more informative to create the decision boundary. Furthermore, the experiments based on University of California, Irvine and Knowledge Extraction based on Evolutionary Learning also show that the proposed method has excellent performances. Even for the minority class in imbalance data sets, the proposed method can also have a good classification.
Dongdong Li 0003, Xinlei Xu, Zhe Wang 0002, Chenjie Cao, Minguang Wang
Int. J. Intell. Syst.1
2022 Gravitation balanced multiple kernel learning for imbalanced classification
Mengping Yang, Zhe Wang 0002, Yanqiong Li, Yangming Zhou, Dongdong Li 0003, Wenli Du
Neural Comput. Appl.5
2022 Geometric imbalanced deep learning with feature scaling and boundary sample mining
Zhe Wang 0002, Qida Dong, Wei Guo 0023, Dongdong Li 0003, Jing Zhang 0041, Wenli Du
Pattern Recognit.4
2022 Learning to Capture the Query Distribution for Few-Shot Learning
abstract
In the Few-Shot Learning (FSL), much of the related efforts only rely on the few available labeled samples (support set) building approach. However, the challenge is that the support set is easy-to-be-biased, so that they cannot be competent prototypes and are hard to represent the class distribution, leading to performance bottlenecks. In this paper, we propose to solve this obstacle by capturing the distribution of the unlabeled samples (query set). We propose two sampling methods: DeepSearch ($\cal DS$) and WideSearch ($\cal WS$). Both approaches are simple to implement and have no trainable parameters. They search the query samples near to the support set in different manners. Afterward, the statistic information is calculated, and we generate the latent samples according to it. The generated latent set is promising. First, it brings the query set distribution information to the classifier, which significantly improves the performance of the cross-entropy-based classifier. Second, it helps the support set become the better prototypes, which boosts the performance of the prototype-based classifier. Third, we find few latent samples are enough to boost the performance. Abundant experiments prove the proposed method achieves state-of-the-art performance on the few-shot tasks. Finally, rich ablation studies explain the compelling details of our approach.
Ziqiu Chi, Zhe Wang 0002, Mengping Yang, Dongdong Li 0003, Wenli Du
IEEE Trans. Circuits Syst. Video Technol.4
2022 Semantic Supplementary Network With Prior Information for Multi-Label Image Classification
abstract
The multi-label image classification problem is one of the most important problems in the field of computer vision, which needs to predict and output all the labels in an image. Multiple labels to be classified in an image increases the difficulty of image classification, and multi-label image classification usually requires additional attention to the positions of the object with different scales and poses. Hence, how to use the dependency relationship between labels to improve the recognition accuracy is an important problem when the object is difficult to directly identify. In this paper, we propose a designed network called the Semantic Supplementary Network with Prior Information (SSNP) to address this problem. The proposed SSNP first generates prior information by using a prior information network with different convolutional layers. Then the semantic supplementary module generates semantic information of the potential labels that is highly relevant to the current information based on the prior information, thereby effectively using the dependency relationship between the labels to improve the classification accuracy. Different from existing methods which pay more attention to the image feature extraction process, we focus on the impact of high-level semantic information generated after feature extraction on the results and tap the potential of high-level semantic information through a semantic supplementary module to strengthen the potential dependence between labels. Experimental results on public benchmark datasets demonstrate that the proposed architecture achieves the state-of-the-art performance, especially when predicting some semantically dependent labels.
Zhe Wang 0002, Zhongli Fang, Dongdong Li 0003, Hai Yang 0002, Wenli Du
IEEE Trans. Circuits Syst. Video Technol.3
2022 Globalized Multiple Balanced Subsets With Collaborative Learning for Imbalanced Data
abstract
The skewed distribution of data brings difficulties to classify minority and majority samples in the imbalanced problem. The balanced bagging randomly undersampes majority samples several times and combines the selected majority samples with minority samples to form several balanced subsets, in which the numbers of minority and majority samples are roughly equal. However, the balanced bagging is the lack of a unified learning framework. Moreover, it fails to concern the connection of all subsets and the global information of the entire data distribution. To this end, this article puts several balanced subsets into an effective learning framework with a criterion function. In the learning framework, one regularization term called$R_{S}$establishes the connection and realizes the collaborative learning of all subsets by requiring the consistent outputs of the minority samples in different subsets. Besides, another regularization term called$R_{W}$provides the global information to each basic classifier by reducing the difference between the direction of the solution vector in each subset and that in the entire dataset. The proposed learning framework is called globalized multiple balanced subsets with collaborative learning (GMBSCL). The experimental results validate the effectiveness of the proposed GMBSCL.
Zonghai Zhu, Zhe Wang 0002, Dongdong Li 0003, Wenli Du
IEEE Trans. Cybern.3
2021 Subtype-GAN: a deep learning approach for integrative cancer subtyping of multi-omics data
abstract
MOTIVATION: The discovery of cancer subtyping can help explore cancer pathogenesis, determine clinical actionability in treatment, and improve patients' survival rates. However, due to the diversity and complexity of multi-omics data, it is still challenging to develop integrated clustering algorithms for tumor molecular subtyping. RESULTS: We propose Subtype-GAN, a deep adversarial learning approach based on the multiple-input multiple-output neural network to model the complex omics data accurately. With the latent variables extracted from the neural network, Subtype-GAN uses consensus clustering and the Gaussian Mixture model to identify tumor samples' molecular subtypes. Compared with other state-of-the-art subtyping approaches, Subtype-GAN achieved outstanding performance on the benchmark datasets consisting of ∼4000 TCGA tumors from 10 types of cancer. We found that on the comparison dataset, the clustering scheme of Subtype-GAN is not always similar to that of the deep learning method AE but is identical to that of NEMO, MCCA, VAE and other excellent approaches. Finally, we applied Subtype-GAN to the BRCA dataset and automatically obtained the number of subtypes and the subtype labels of 1031 BRCA tumors. Through the detailed analysis, we found that the identified subtypes are clinically meaningful and show distinct patterns in the feature space, demonstrating the practicality of Subtype-GAN. AVAILABILITYAND IMPLEMENTATION: The source codes, the clustering results of Subtype-GAN across the benchmark datasets are available at https://github.com/haiyang1986/Subtype-GAN. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Hai Yang 0002, Rui Chen 0021, Dongdong Li 0003, Zhe Wang 0002
Bioinform.3
2021 Multi-kernel Support Vector Data Description with boundary information
Wei Guo 0023, Zhe Wang 0002, Sisi Hong, Dongdong Li 0003, Hai Yang 0002, Wen Du
Eng. Appl. Artif. Intell.4
2021 Speech emotion recognition using recurrent neural networks with directional self-attention
Dongdong Li 0003, Jinlin Liu, Linyu Sun, Zhe Wang 0002
Expert Syst. Appl.1
2021 Entropy-based hybrid sampling ensemble learning for imbalanced data
abstract
Sampling method is one of the most commonly used techniques in dealing with imbalanced data. Most of the existing undersampling methods randomly select samples from negative class with replacement. However, it may lose some important information of the training data. Moreover, increasing the positive data by oversampling in high imbalanced situations may cause the overlapping problem. To overcome these problems, this paper proposes a hybrid sampling method. The method takes the distributions of the training data into consideration by the information entropy, thus distinguishing the important samples in the undersampling procedure. Meanwhile, since the positive data only extend to the size of each subset of the negative class in the oversampling, the overlapping problem is relieved. Further, the method retains all the data in the training procedure and generates various data views from the original training data. Then each view is handled with an individual basic classifier. Finally, all the basic classifiers are combined by the ensemble method. The newly proposed method is named as Entropy-based Hybrid Sampling Ensemble Learning (EHSEL). In addition, the EHSEL is applied to three different kinds of basic classifiers to validate its robustness. Experiments results show the great effectiveness of the EHSEL on real-world imbalanced data sets.
Dongdong Li 0003, Ziqiu Chi, Bolu Wang, Zhe Wang 0002, Hai Yang 0002, Wenli Du
Int. J. Intell. Syst.1
2021 Exploiting the potentialities of features for speech emotion recognition
Dongdong Li 0003, Zhe Wang 0002, Daqi Gao
Inf. Sci.1
2021 BLSTM and CNN Stacking Architecture for Speech Emotion Recognition
Dongdong Li 0003, Linyu Sun, Xinlei Xu, Zhe Wang 0002, Jing Zhang 0041, Wenli Du
Neural Process. Lett.1
2021 FLDNet: Frame-Level Distilling Neural Network for EEG Emotion Recognition
abstract
Based on the current research on EEG emotion recognition, there are some limitations, such as hand-engineered features, redundant and meaningless signal frames and the loss of frame-to-frame correlation. In this paper, a novel deep learning framework is proposed, named the frame-level distilling neural network (FLDNet), for learning distilled features from the correlations of different frames. A layer named the frame gate is designed to integrate weighted semantic information on multiple frames to remove redundant and meaningless signal frames. A triple-net structure is introduced to distill the learned features net by net to replace the hand-engineered features with professional knowledge. Specifically, one neural network is normally trained for several epochs. Then, a second network of the same structure will be initialized again to learn the extracted features from the frame gate of the first neural network based on the output of the first net. Similarly, the third net improves the features based on the frame gate of the second network. To utilize the representation ability of the triple neural network, an ensemble layer is conducted to integrate the discriminative ability of the proposed framework for final decisions. Consequently, the proposed FLDNet provides an effective method for capturing the correlation between different frames and automatically learn distilled high-level features for emotion recognition. The experiments are carried out in a subject-independent emotion recognition task on public emotion datasets of DEAP and DREAMER benchmarks, which have demonstrated the effectiveness and robustness of the proposed FLDNet.
Zhe Wang 0002, Tianhao Gu, Dongdong Li 0003, Hai Yang 0002, Wenli Du
IEEE J. Biomed. Health Informatics4
2020 Multiple Random Empirical Kernel Learning with Margin Reinforcement for imbalance problems
Zhe Wang 0002, Lilong Chen, Dongdong Li 0003, Daqi Gao
Eng. Appl. Artif. Intell.4
2020 Multiple Universum Empirical Kernel Learning
Zhe Wang 0002, Sisi Hong, Lijuan Yao, Dongdong Li 0003, Wenli Du, Jing Zhang 0041
Eng. Appl. Artif. Intell.4
2020 Weight-based multiple empirical kernel learning with neighbor discriminant constraint for heart failure mortality prediction
Zhe Wang 0002, Bolu Wang, Yangming Zhou, Dongdong Li 0003, Yichao Yin
J. Biomed. Informatics4
2020 Entropy and gravitation based dynamic radius nearest neighbor classification for imbalanced problem
Zhe Wang 0002, Yanqiong Li, Dongdong Li 0003, Zonghai Zhu, Wenli Du
Knowl. Based Syst.3
2020 Object semantics sentiment correlation analysis enhanced image sentiment classification
Jing Zhang 0041, Dongdong Li 0003, Zhe Wang 0002
Knowl. Based Syst.4
2020 NearCount: Selecting critical instances based on the cited counts of nearest neighbors
Zonghai Zhu, Zhe Wang 0002, Dongdong Li 0003, Wenli Du
Knowl. Based Syst.3
2020 Multi-matrices entropy discriminant ensemble learning for imbalanced problem
Zhe Wang 0002, Zhaozhi Chen, Jing Zhang 0041, Wenli Du, Dongdong Li 0003
Neural Comput. Appl.6
2020 Efficient matrixized classification learning with separated solution process
Zonghai Zhu, Zhe Wang 0002, Dongdong Li 0003, Wenli Du, Jing Zhang 0041
Neural Comput. Appl.3
2020 Multiple Partial Empirical Kernel Learning with Instance Weighting and Boundary Fitting
Zonghai Zhu, Zhe Wang 0002, Dongdong Li 0003, Wenli Du, Yangming Zhou
Neural Networks3
2020 Cancer classification based on chromatin accessibility profiles with deep adversarial learning model
abstract
Given the complexity and diversity of the cancer genomics profiles, it is challenging to identify distinct clusters from different cancer types. Numerous analyses have been conducted for this propose. Still, the methods they used always do not directly support the high-dimensional omics data across the whole genome (Such as ATAC-seq profiles). In this study, based on the deep adversarial learning, we present an end-to-end approach ClusterATAC to leverage high-dimensional features and explore the classification results. On the ATAC-seq dataset and RNA-seq dataset, ClusterATAC has achieved excellent performance. Since ATAC-seq data plays a crucial role in the study of the effects of non-coding regions on the molecular classification of cancers, we explore the clustering solution obtained by ClusterATAC on the pan-cancer ATAC dataset. In this solution, more than 70% of the clustering are single-tumor-type-dominant, and the vast majority of the remaining clusters are associated with similar tumor types. We explore the representative non-coding loci and their linked genes of each cluster and verify some results by the literature search. These results suggest that a large number of non-coding loci affect the development and progression of cancer through its linked genes, which can potentially advance cancer diagnosis and therapy.
Hai Yang 0002, Dongdong Li 0003, Zhe Wang 0002
PLoS Comput. Biol.3
2020 Collaborative and geometric multi-kernel learning for multi-class classification
Zhe Wang 0002, Zonghai Zhu, Dongdong Li 0003
Pattern Recognit.3
2020 Geometric Structural Ensemble Learning for Imbalanced Problems
abstract
The classification on imbalanced data sets is a great challenge in machine learning. In this paper, a geometric structural ensemble (GSE) learning framework is proposed to address the issue. It is known that the traditional ensemble methods train and combine a series of basic classifiers according to various weights, which might lack the geometric meaning. Oppositely, the GSE partitions and eliminates redundant majority samples by generating hyper-sphere through the Euclidean metric and learns basic classifiers to enclose the minority samples, which achieves higher efficiency in the training process and seems easier to understand. In detail, the current weak classifier builds boundaries between the majority and the minority samples and removes the former. Then, the remaining samples are used to train the next. When the training process is done, all of the majority samples could be cleaned and the combination of all basic classifiers is obtained. To further improve the generalization, two relaxation techniques are proposed. Theoretically, the computational complexity of GSE could approach O(nd log(nmin) log(nmaj)). The comprehensive experiments validate both the effectiveness and efficiency of GSE.
Zonghai Zhu, Zhe Wang 0002, Dongdong Li 0003, Yujin Zhu, Wenli Du
IEEE Trans. Cybern.3
2019 Multiple Empirical Kernel Learning with Discriminant Locality Preservation
abstract
Multiple Kernel Learning (MKL) algorithm effectively combines different kernels to improve the performance of classification. Most MKL algorithms implicitly map samples into feature space by the form of inner-product. In contrast, Multiple Empirical Kernel Learning (MEKL) can explicitly map the input spaces into feature spaces so that the mapped feature vectors are explicitly represented, which is easy to process and analyze the adaptability of kernels for input space. Meanwhile, in order to pay attention to the structure and discriminant information of samples in empirical feature space, inspired by discriminant locality preserving projections, we introduce the discriminant locality preservation regularization into MEKL framework to propose the Multiple Empirical Kernel Learning with Discriminant Locality Preservation (MEKL-DLP). Experiments conducted on real-world datasets validate the effectiveness of the proposed MEKL-DLP compared with the classical kernel-based algorithms and state-of-art MKL algorithms.
Bolu Wang, Dongdong Li 0003, Zhe Wang 0002
ACML2
2019 Web image annotation based on Tri-relational Graph and semantic context analysis
Jing Zhang 0041, Ti Tao, Yakun Mu, Dongdong Li 0003, Zhe Wang 0002
Eng. Appl. Artif. Intell.5
2019 Cost-sensitive Fuzzy Multiple Kernel Learning for imbalanced problem
Zhe Wang 0002, Bolu Wang, Dongdong Li 0003, Jing Zhang 0041
Neurocomputing4
2019 Tree-based space partition and merging ensemble learning framework for imbalanced problems
Zonghai Zhu, Zhe Wang 0002, Dongdong Li 0003, Wenli Du
Inf. Sci.3
2017 GMFLLM: A general manifold framework unifying three classic models for dimensionality reduction
Yujin Zhu, Zhe Wang 0002, Daqi Gao, Dongdong Li 0003
Eng. Appl. Artif. Intell.4
2017 Entropy-based fuzzy support vector machine for imbalanced datasets
Zhe Wang 0002, Dongdong Li 0003, Daqi Gao, Hongyuan Zha
Knowl. Based Syst.3
2017 Locality sensitive discriminant matrixized learning machine
Zhe Wang 0002, Dongdong Li 0003, Yujin Zhu, Chenjie Cao
Knowl. Based Syst.3
2017 Regularized Matrix-Pattern-Oriented Classification Machine with Universum
Dongdong Li 0003, Yujin Zhu, Zhe Wang 0002, Chuanyu Chong, Daqi Gao
Neural Process. Lett.1
2015 MPEKDyL: Efficient multi-partial empirical kernel dynamic learning
Zhe Wang 0002, Daqi Gao, Dongdong Li 0003
Knowl. Based Syst.4
2015 Affect-insensitive speaker recognition systems via emotional speech clustering using prosodic features
Dongdong Li 0003, Yubo Yuan 0001, Zhaohui Wu 0001, Yingchun Yang
Neural Comput. Appl.1
2007 Affect-Insensitive Speaker Recognition by Feature Variety Training
Dongdong Li 0003, Yingchun Yang
ACII1
2006 Rules Based Feature Modification for Affective Speaker Recognition
abstract
One of the largest challenges in speaker recognition applications is dealing with speaker-emotion variability. In this paper, we further investigate the rules based feature modification for robust speaker recognition with emotional speech. Specifically, we learn the rules of prosodic features modification from a small amount of the content matched source-target pairs. Features with emotion information are adapted from the prevalent neutral features by applying the modification rules. The converted features are trained together with the neutral features to build the speaker models. The effects of individual and combined modifications of duration, pitch and amplitude are also studied using EPST dataset recorded by 8 professional actors with 14 kinds of emotion expressiveness. It demonstrates that duration modifications play the most important role; and that, pitch modifications are more effective than amplitude modifications. Promising result with an improved identification rate by 7.83% is achieved compared to the traditional speaker recognition
Zhaohui Wu 0001, Dongdong Li 0003, Yingchun Yang
ICASSP (1)2
2005 Emotion-State Conversion for Speaker Recognition
Dongdong Li 0003, Yingchun Yang, Zhaohui Wu 0001
ACII1
2005 Combining voiceprint and face biometrics for speaker identification using SDWS
Dongdong Li 0003, Yingchun Yang, Zhaohui Wu 0001
INTERSPEECH1