VLDB 2026 Research / reviewers in the wild / expert
Zhaohong Deng
dblp:63/6198
· DBLP profile ↗
119ranked-venue papers
18as first author
68since 2021 · last 2027
0000-0002-0790-6492ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 73 · 13 first-author · 34 since 2021Applied, interdisciplinary, general and emerging computing · 26 · 23 since 2021Databases, data management, data science and information retrieval · 11 · 5 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 6 since 2021Human-computer interaction and ubiquitous computing · 5 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Evidential gating mixture-of-experts based on spatial and temporal transformers for uncertainty-aware engagement estimation
Zhengrui Wang, Hongqiang Shen, Zhaohong Deng |
Expert Syst. Appl. | 4 |
| 2026 | PatchET: Learning Enzyme Temperature Properties Through Patch-Based Neural ArchitecturesabstractUnderstanding enzyme thermal properties is essential for biotechnology and protein engineering, yet experimental measurements of attributes such as temperature optimum, stability, and range remain labor-intensive and costly. Prior studies have shown that specific regions within enzyme sequences disproportionately influence thermal behavior—an aspect often overlooked by existing deep learning models. In this work, we introduce PatchET, a biologically inspired deep learning model that predicts enzyme thermal properties directly from amino acid sequences. PatchET employs a dual-stage, patch-based architecture that captures both intra-patch local features and inter-patch global dependencies, reflecting the hierarchical nature of protein thermal adaptation. Alongside the model, we curate a comprehensive benchmark, including a refined dataset for temperature optimum and the first publicly available dataset for temperature range prediction. PatchET achieves state-of-the-art performance across three key tasks—temperature optimum, stability, and range—and serves as the first dedicated model for temperature range prediction. Extensive ablation studies further validate the effectiveness of our architectural design. Together, PatchET and the accompanying benchmark provide a unified and generalizable framework for modeling enzyme thermal properties, offering new tools for the rational design of thermostable enzymes. Longbing Cao, Zhaohong Deng |
AAAI | 4 |
| 2026 | Ensemble clustering method via learning enhanced consensus adjacency matrices
Zekang Bian, Jinwei Sun, Qidong Dai, Qiongdan Lou, Zhaohong Deng, Shitong Wang 0001 |
Neurocomputing | 6 |
| 2026 | Deep learning for multi-step financial time series forecasting with spectral energy-guided variational mode decomposition
Zhanhua Dong, Zijian Xu 0016, Qiujun Pan, Zhaohong Deng, Jianlei Liu |
Neurocomputing | 4 |
| 2026 | A novel TSK fuzzy system incorporating multi-view collaborative transfer learning for personalized epileptic EEG detection
Andong Li, Zhaohong Deng, Qiongdan Lou |
Neurocomputing | 2 |
| 2026 | FKSUDDAPre: A drug-disease association prediction framework based on F-TEST feature selection and AMDKSU resampling with interpretability analysisabstractIn drug discovery and therapeutic research, the prediction of drug-disease associations (DDAs) holds significant scientific and clinical value. Drug molecules exert their effects by precisely identifying disease-related biological targets, systematically modulating the entire pharmacological process from absorption, distribution, and metabolism to final efficacy. Accurate prediction of drug-disease associations not only facilitates an in-depth understanding of molecular mechanisms of drug action but also provides critical theoretical foundations for drug repositioning and personalized medicine. While traditional prediction methods based on in vitro experiments and clinical statistics yield reliable results, they suffer from inherent drawbacks such as long development cycles, substantial resource consumption, and low throughput. In contrast, emerging machine learning techniques offer a promising solution to these bottlenecks, enabling the intelligent and efficient discovery of potential drug-disease association networks and significantly improving drug development efficiency. However, it is noteworthy that existing machine learning methods still face significant challenges in practical applications: the complexity of feature construction raises the threshold for data processing; data sparsity constrains the depth of information mining; and the pervasive issue of sample imbalance poses a severe challenge to the model's predictive accuracy and generalization performance. In this study, we developed an efficient and accurate framework for drug-disease association prediction named FKSUDDAPre. The model employs a multi-modal feature fusion strategy: on one hand, it leverages an ensemble of Mol2vec and K- BERT to deeply capture the semantic features of drug molecular fingerprints; on the other hand, it integrates Medical Subject Headings (MeSH) with DeepWalk to effectively reduce the dimensionality of disease features while preserving their relational structure. To address the class imbalance problem, FKSUDDAPre designed an optimization algorithm called AMDKSU, which combined clustering with an improved distance metric strategy, significantly enhancing the discriminative power of the sample set. For data processing, F-test was employed for feature importance ranking, effectively reducing data dimensionality and improving model generalization. For the predictive architecture, FKSUDDAPre proposed a novel ensemble framework composed of XGBoost, Decision Tree, Random Forest, and HyperFast. By employing a dynamic weight allocation strategy, this ensemble effectively harnesses the complementary strengths of these models to achieve significantly enhanced predictive performance. Rigorous validation demonstrated the system's outstanding performance across multiple evaluation metrics, with an average AUC of 0.9725, improving the AUC by approximately 3.88% compared to the best-performing baseline model. In the prediction of Alzheimer's disease and Parkinson's disease, 80% and 60% of the top 10 candidate drugs recommended by FKSUDDAPre, respectively, had been confirmed by literature, demonstrating the model's good practical application potential. Furthermore, we conducted a LIME-based feature importance analysis on the model's predictions, visualizing the correlations between features and the target variable to demonstrate the model's interpretability. A cross-platform, user-friendly visualization tool had also been developed using the PyQt5 framework. Yun Zuo 0001, Ge Hua, Xiangrong Liu, Xiangxiang Zeng, Zhaohong Deng |
PLoS Comput. Biol. | 7 |
| 2026 | Colonic polyp segmentation based on transformer-convolutional neural networks fusion
Chenxi Luo, Zhaohong Deng, Qiongdan Lou, Zhuangzhuang Zhao, Yuxi Ge, Shudong Hu |
Pattern Recognit. | 3 |
| 2026 | MTDNet: A crowd counting network based on a multiscale transformer and dilated convolution
Chongle Peng, Qingbing Sang, Xiaojun Wu 0001, Zhaohong Deng, Lixiong Liu |
Signal Process. Image Commun. | 4 |
| 2026 | SMENET: A Multi-View Semantic Model for Multi-Level Enzyme Function PredictionabstractComprehending biological reproduction and cellular metabolism is facilitated by the Enzyme Commission, which matches protein sequences to the biochemical reactions they catalyse through EC numbers. In recent years, several methods have been proposed for predicting enzyme function. However, these methods still encounter challenges. Firstly, traditional methods for manually designing enzyme features are complex and cumbersome, lacking an effective generalized method for embedding enzyme sequences. Secondly, the distribution gap between different enzymes is significant, which resulting in existing methods struggling to predict multilevel enzyme functions. Thirdly, traditional enzyme function prediction models only extract single view feature of enzyme, so there is still room for further improving the ability of these models to extract enzyme data. To address these challenges, a new multilevel enzyme function prediction model (SMENET) based on multi-view semantics is proposed. This method uses protein large language model to extract semantic information. Subsequently, this semantic information is fed into multiple information extraction network modules, followed by using Biologic Sematic Attention to integrate these views' information. Finally, a multi-view adaptive fusion network is designed to extract the best common representation between multiple semantic views. Extensive experiments were conducted on multiple datasets to validate the effectiveness of SMENET. Hanwen Zhou, Wei Zhang 0221, Zhaohong Deng, Guanjin Wang, Zhisheng Wei, Xiaoyong Pan, Hong-Bin Shen, Dongjun Yu, Jing Wu 0030 |
IEEE Trans. Comput. Biol. Bioinform. | 3 |
| 2026 | Fuzzy Ensemble Clustering Method via Learning Enhanced Fuzzy Connective Matrices
Zekang Bian, Qidong Dai, Te Zhang, Zhaohong Deng, Shitong Wang 0001 |
IEEE Trans. Fuzzy Syst. | 7 |
| 2026 | GFS-Edge: Graph Fuzzy System for Edge PredictionabstractEdge prediction (EP) is a long-standing and fundamental task in graph-structured data analysis, and has served as a key driver in the advancement of graph learning. Traditional EP methods can be broadly classified into three major categories: heuristic approaches, embedding-based methods, and graph neural network (GNN)-based techniques. While each category has achieved notable performance improvements, these methods often fall short in handling uncertainty and generally lack robust interpretability. These limitations highlight an urgent need for novel approaches that can effectively address uncertain EP tasks while offering strong interpretability and rich feature representations. To tackle these challenges, we propose a novel graph fuzzy system (GFS) for edge prediction, termed GFS-edge, and systematically introduces its core concepts, model architecture, and construction methodology. Specifically, we first define key concepts of GFS-edge, including rule bases, fuzzy sets, and edge consequent processing module (ECPM). We then present a general modeling framework for GFS-edge, detailing the formulation of rule antecedents and consequents. Subsequently, a learning framework is introduced, in which edge-based clustering is employed for antecedent generation, while a linear variational graph auto-encoder (LVGAE) is adopted as the ECPM to facilitate consequent learning. Extensive experiments are conducted on 12 benchmark datasets to evaluate the performance of GFS-edge. The results consistently demonstrate that GFS-edge outperforms existing methods from all three traditional EP categories, showcasing superior performance while effectively addressing uncertainty modeling and interpretability challenges. Fuping Hu, Zhaohong Deng, Kup-Sze Choi, Shitong Wang 0001 |
IEEE Trans. Fuzzy Syst. | 2 |
| 2026 | Enhanced One-Step Incomplete Multiview Fuzzy Clustering With Dual Representation LearningabstractMulti-view fuzzy clustering has attracted increasing attention owing to its strong clustering performance and inherent ability to effectively model uncertainty. However, most existing methods rely on the unrealistic assumption that all views are fully observed, which rarely holds in practice. Although several methods have been proposed to address incomplete multi-view data, they typically focus only on extracting shared information across views while overlooking view-specific information. Moreover, they often tend to neglect missing views imputation, a key mechanism for handling incomplete data. Furthermore, by separating representation learning from clustering, many existing frameworks yield representations that are not necessarily optimal for clustering, thus compromising robustness. To address these limitations and based on fuzzy clustering, a novel enhanced one step incomplete multi-view clustering method (IMVFCM_DRL) is proposed in this paper. First, to effectively handle incomplete multi-view data, we construct a new representation learning framework that explicitly integrates missing-view imputation. Second, to fully exploit the multi-view information, a dual information learning strategy is introduced to jointly capture both common and view-specific information. Finally, a unified one-step fuzzy clustering framework with weighted structure preservation is developed, ensuring that representation learning and fuzzy clustering are jointly optimized. The experiments conducted on various multi-view datasets demonstrate the effectiveness of IMVFCM_DRL. The codes are available at https://github.com/BBKing49/IMVFCM_DRL. Wei Zhang 0221, Zhaohong Deng, Weiping Ding 0001, Jun Zhou 0029, Te Zhang, Kup-Sze Choi, Shitong Wang 0001 |
IEEE Trans. Fuzzy Syst. | 2 |
| 2026 | Fuzzy Rule-Guided Multiview Differentiable Representation Learning With Dual-Space Information ExtractionabstractEffectively extracting discriminative information from multi-view data remains a key challenge in multi-view learning. Existing methods typically focus on exploring inter-view consistency via linear or nonlinear transformation. While nonlinear methods tend to yield better performance, their limited transparency and interpretability limit their practical applications. Takagi-Sugeno-Kang Fuzzy Systems (TSK-FS), as a rule-based model with high interpretability, have been applied to multi-view tasks. However, prior methods either rely solely on antecedent components for nonlinear modeling or integrate deep neural networks into the consequent part, thereby compromising model interpretability. To address these challenges, we propose Fuzzy Rule-guided Multi-view Differentiable Representation Learning (FRMVDRL). Specifically, in our framework, antecedent parameters of TSK-FS are first used to map data into highdimensional fuzzy space. Then during the learning of consequent parameters, a dual information extraction mechanism is proposed to jointly capture shared knowledge across views and view-specific knowledge. Moreover, a second-order geometric structure preservation mechanism is constructed to exploit structural information at both the instance and the instance-pair level. To enhance discriminability of the learned representations, a biorthogonal constraint alongside a Shannon entropy mechanism is introduced. Finally, to balance model performance and interpretability, we introduce a novel multi-view differentiable optimization strategy that incorporates learnable parameters to expand the solution space while preserving the structure of traditional optimization. Extensive experiments on benchmark multi-view datasets demonstrate the effectiveness of the FRMVDRL. Wei Zhang 0221, Jun Zhou 0029, Guanjin Wang, Zhaohong Deng, Weiping Ding 0001, Witold Pedrycz |
IEEE Trans. Fuzzy Syst. | 4 |
| 2026 | Generative Fuzzy System for Sequence-to-Sequence Learning via Rule-Based InferenceabstractGenerative models (GMs), particularly large language models (LLMs), have garnered significant attention in machine learning and artificial intelligence for their ability to generate new data by learning the statistical properties of training data and creating data that resemble the original data. This capability offers a wide range of applications across various domains. However, the complex structures and numerous model parameters of GMs obscure the input-output processes and complicate the understanding and control of the outputs. Moreover, the purely data-driven learning mechanism limits GMs' abilities to acquire broader knowledge. There remains substantial potential for enhancing the robustness and generalization capabilities of GMs. In this work, we leverage fuzzy system, a classical modeling method, to combine both data-driven and knowledge-driven mechanisms for generative tasks. We propose a novel generative fuzzy system framework, named GenFS, which integrates the deep learning capabilities of GMs with the term-based interpretability and dual-driven mechanisms of fuzzy systems. Specifically, we propose an end-to-end GenFS-based model for sequence generation, called FuzzyS2S. A series of test studies were conducted on 12 datasets, covering three distinct categories of generative tasks: machine translation, code generation, and summary generation. The results demonstrate that FuzzyS2S outperforms the transformer in terms of accuracy and fluency. Furthermore, it exhibits better performance than state-of-the-art models T5 and CodeT5 for some application scenarios. Hailong Yang 0001, Zhaohong Deng, Wei Zhang 0221, Zhuangzhuang Zhao, Guanjin Wang, Kup-Sze Choi |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2026 | Fuzzy Rule-Based Differentiable Representation LearningabstractRepresentation learning is a key area in machine learning and deep learning, focusing on extracting meaningful features to support downstream tasks such as classification and clustering. Current mainstream representation learning methods primarily rely on nonlinear data mining techniques such as kernel methods and deep neural networks (DNNs) to extract abstract knowledge from complex datasets. However, most of them are "black-box" methods, lacking transparency and interpretability in the learning process, which constrain their practical utility. To this end, this article introduces a novel representation learning method called fuzzy rule-based differentiable representation learning (FRDRL), which is grounded in an interpretable fuzzy rule-based model. Specifically, it is built upon the Takagi-Sugeno-Kang fuzzy system (TSK-FS) to map input data to a high-dimensional fuzzy feature space through the antecedent part of the TSK-FS. Subsequently, a novel differentiable optimization method is proposed for learning in the consequent part, which preserves interpretability and transparency while effectively capturing nonlinear relationships in the data. By retaining the essence of traditional optimization and parameterizing key components as differentiable modules, the method improves performance without sacrificing interpretability. Moreover, a second-order geometry preservation strategy is incorporated to further improve robustness. Extensive evaluations conducted on various benchmark datasets validate the superiority of the proposed method. The source codes are available at https://github.com/BBKing49/FEDRL. Wei Zhang 0221, Zhaohong Deng, Guanjin Wang, Kup-Sze Choi |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2025 | ADKcat: Enhanced Enzyme Turnover Number Prediction with Adaptive Data Augmentation and Dual Information ExplorationabstractThe enzyme turnover number is a key metric for catalytic efficiency. While recent deep learning models have integrated enzyme and substrate modal information for predicting turnover numbers, the following challenges remain. First, existing datasets are limited in scale, inconsistent, and lack standardization. Second, current models emphasize consistency information across modalities while ignoring the specific information within each modality. Third, the imbalanced distribution of measured turnover numbers leads to poor performance of existing methods in extreme value ranges. To address these challenges, we first construct Kinetic-DB, a large, standardized dataset compiled from public sources. Based on this, we propose a novel method, ADKcat, for enzyme turnover number prediction. Specifically, to mitigate the challenge of imbalanced data, an adaptive data augmentation module is first constructed to enrich both enzyme and substrate sequences. Then, two pretrained language models-ESM-2 for enzymes and Mole-BERT for substrates-are used for embedding extraction. Furthermore, to fully explore common and specific information, we introduce a dual information exploration module and enhance it with domain classification and distribution alignment loss functions. Finally, an adaptive density weighted network is further applied to improve prediction accuracy and robustness. Experiments show that ADKcat outperforms state-of-the-art methods, especially in extreme turnover ranges. This makes it a promising tool for enzyme engineering, drug discovery, and synthetic biology. Weiping Ding 0001, Wei Zhang 0221, Zhaohong Deng |
BIBM | 5 |
| 2025 | Self-Supervised Learning and Image-Prompt Fusion for AIGC Image Quality AssessmentabstractWith the rapid advancement of artificial intelligence, the field of Artificial Intelligence Generated Content (AIGC) has seen significant growth. As AI-generated images (AIGIs) become increasingly prevalent, the AIGC image quality assessment(AIGCIQA) has gained critical importance. However, traditional image quality assessment methods struggle to account for the complex relationship between generated images and their corresponding text prompts, leading to suboptimal performance in AIGCIQA tasks. Additionally, the scarcity of subjective annotation data within AIGCIQA datasets limits the effectiveness of deep learning models. To address these challenges, we propose a novel self-supervised task that utilizes intermediate image sequences generated during Text-to-Image (T2I) model generation as the pre-training data. We use the number of model iterations as pseudo-labels based on the positive correlation between image quality and the number of model iterations. Due to the relative coarseness of this pseudo-label as a supervised signal, we also introduce a linear interpolation method for optimization. Additionally, we designed a framework that effectively fuses image and text features. The framework considers the dynamic fading properties of intermediate image sequences similar to video clips. It utilizes a specialized text prompt template and extracts image features through a cross-attention mechanism, which significantly improves model performance. Experimental results show that the proposed method achieves state-of-the-art performance on three mainstream AIGCIQA datasets. Qingbing Sang, Zhaohong Deng, Xiaojun Wu 0001 |
ICASSP | 3 |
| 2025 | Dual-Branch Dynamic Coupling Weakly Supervised Learning for Class-Incremental Histopathological Region Segmentation
Xiaoyan Hong, Jiansong Fan, Zhaohong Deng |
MICCAI (10) | 3 |
| 2025 | MlyPredCSED: based on extreme point deviation compensated clustering combined with cross-scale convolutional neural networks to predict multiple lysine sites in humanabstractIn post-translational modification, covalent bonds on lysine and attached chemical groups significantly change proteins' physical and chemical properties. They shape protein structures, enhance function and stability, and are vital for physiological processes, affecting health and disease through mechanisms like gene expression, signal transduction, protein degradation, and cell metabolism. Although lysine (K) modification sites are considered among the most common types of post-translational modifications in proteins, research on K-PTMs has largely overlooked the synergistic effects between different modifications and lacked the techniques to address the problem of sample imbalance. Based on this, the Extreme Point Deviation Compensated Clustering (EPDCC) Undersampling algorithm was proposed in this study and combined with Cross-Scale Convolutional Neural Networks (CSCNNs) to develop a novel computational tool, MlyPredCSED, for simultaneously predicting multiple lysine modification sites. MlyPredCSED employs Multi-Label Position-Specific Triad Amino Acid Propensity and the physicochemical properties of amino acids to enhance the richness of sequence information. To address the challenge of sample imbalance, the innovative EPDCC Undersampling technique was introduced to adjust the majority class samples. The model's training and testing phase relies on the advanced CSCNN framework. MlyPredCSED, through cross-validation and testing, outperformed existing models, especially in complex categories with multiple modification sites. This research not only provides an efficient method for the identification of lysine modification sites but also demonstrates its value in biological research and drug development. To facilitate efficient use of MlyPredCSED by researchers, we have specifically developed an accessible free web tool: http://www.mlypredcsed.com. Yun Zuo 0001, Xingze Fang, Jiankang Chen, Jiayi Ji, Xiangrong Liu, Xiangxiang Zeng, Zhaohong Deng, Hongwei Yin, Anjing Zhao |
Briefings Bioinform. | 9 |
| 2025 | m2ST: dual multi-scale graph clustering for spatially resolved transcriptomicsabstractMOTIVATION: Spatial clustering is a key analytical technique for exploring spatial transcriptomics data. Recent graph neural network-based methods have shown promise in spatial clustering but face notable challenges. One significant issue is that analyzing the functions and complex mechanisms of organisms from a single scale is difficult and most methods focus exclusively on the single-scale representation of transcriptomic data, potentially limiting the discriminative power of extracted features for spatial domain clustering. Furthermore, classical clustering algorithms are often applied directly to latent representation, making it a worthwhile endeavor to explore a tailored clustering method to further improve the accuracy of spatial domain annotation. RESULTS: To address these limitations, we propose m2ST, a novel dual multi-scale graph clustering method. m2ST first uses a multi-scale masked graph autoencoder to extract representations across different scales from spatial transcriptomic data. To effectively compress and distill meaningful knowledge embedded in the data, m2ST introduces a random masking mechanism for node features and uses a scaled cosine error as the loss function. Additionally, we introduce a tailored multi-scale clustering framework that integrates scale-common and scale-specific information exploration into the clustering process, achieving more robust annotation performance. Shannon entropy is finally utilized to dynamically adjust the importance of different scales. Extensive experiments on multiple spatial transcriptomic datasets demonstrate the superior performance of m2ST compared to existing methods. AVAILABILITY AND IMPLEMENTATION: https://github.com/BBKing49/m2ST. Wei Zhang 0221, Hailong Yang 0001, Te Zhang, Zhaohong Deng, Xiaoyong Pan, Hong-Bin Shen, Dongjun Yu, Shitong Wang 0001 |
Bioinform. | 7 |
| 2025 | Food image segmentation based on deep and shallow dual-branch network
Zhiyong Xiao 0001, Zhaohong Deng |
Multim. Syst. | 3 |
| 2025 | HyperACP: A cutting-edge hybrid framework for anticancer peptide classification via scalable feature extraction and adaptive neighbor-based synthesisabstractCancer remains a major contributor to global mortality, constituting a significant and escalating threat to human health. Anticancer peptides (ACPs) have emerged as promising therapeutic agents due to their specific mechanisms of action, pronounced tumor-targeting capability, and low toxicity. Nevertheless, traditional approaches for ACP identification are constrained by their reliance on shallow, hand-crafted sequence features, which fail to capture deeper semantic and structural characteristics. Moreover, such models exhibit limited robustness and interpretability when confronted with practical challenges such as severe class imbalance. To address these limitations, this study proposes HyperACP, an innovative framework for ACP recognition that integrates deep representation learning, adaptive sampling, and mechanistic interpretability. The framework leverages the ESMC protein language model to extract comprehensive sequence features and employs a novel adaptive algorithm, ANBS, to mitigate class imbalance at the decision boundary. For enhanced model transparency, SHAP-Res is incorporated to elucidate the contributions of individual residues to the final predictions. Comprehensive evaluations demonstrate that HyperACP consistently outperforms state-of-the-art methods across multiple datasets and validation protocols-including 10-fold cross-validation and independent test sets-according to metrics such as Accuracy (ACC), Sensitivity (SN), Specificity (SP), Matthews Correlation Coefficient (MCC), and Area Under the Curve (AUC). Furthermore, the model yields biologically interpretable results, pinpointing key residues (K, L, F, G) known to play pivotal roles in anticancer activity. These findings provide not only a robust predictive tool (available at www.hyperacp.com) but also novel insights into the structure-function relationships underlying ACPs. Bangyi Zhang, Yun Zuo 0001, Jiayue Liu, Xiangrong Liu, Xiangxiang Zeng, Zhaohong Deng |
PLoS Comput. Biol. | 7 |
| 2025 | CATransUnetLBP: Accurate Prediction of Protein-Ligand Binding Pockets Using a Hybrid NetworkabstractThe development of intelligent methods capable of predicting protein-ligand binding sites has become a popular research field. Recently, deep learning based methods have been proposed as a promising solution for this task. However, some limitations still exist. For example, the network structure is not optimized for predicting protein binding pockets, which limits the model's capabilities. To address the aforementioned challenges, a novel method called CATransUnetLPB is proposed, in which a new network structure named CATransUnet is designed. The proposed CATransUnet combines CNN and Transformer models to accurately segment binding pocket regions from protein 3D structures. It outperforms existing representative methods on three test sets, demonstrating the effectiveness of optimizing the deep network model for detecting protein ligand binding pockets. Furthermore, we conduct thorough analysis on applying data augmentation to protein data structure and confirm that such technique can enhance the model's generalization ability, thereby ensuring good performance on new protein structures. Moreover, experiments show that the predicted binding pockets from our model can complement the results obtained from other methods. This suggests that integrating our method with existing approaches could further improve the prediction of protein-ligand binding pockets. Cheng Cai, Zhaohong Deng, Andong Li, Yun Zuo 0001, Haoran Chen 0003, Zhisheng Wei, Xiaoyong Pan, Hong-Bin Shen, Dongjun Yu |
IEEE Trans. Comput. Biol. Bioinform. | 2 |
| 2025 | DMMAFS: Protein Function Prediction Based on Multi-Modal Multi-Attention Fusion FeaturesabstractIntelligent prediction of protein function is more efficient and less resource-consuming and has achieved significant progress in recent years. However, most of the current methods are performed solely based on the sequence information of proteins. These methods overlook information of other modalities that the proteins themselves possess, which makes it difficult to achieve the desired predicted results. Furthermore, a few existing methods based on multiple modal information fuse them in a simple splicing manner and fail to fully exploit the complementary relation between different modalities. To address the above-mentioned challenges, we propose Multi-modal Multi-attention fusion Features (DMMAFS), a method based on deep learning, to predict protein function. On the one hand, DMMAFS gains the semantic information embedded in the sequence itself through the self-attention learning of the sequence. On the other hand, DMMAFS employs the 3D structural information of proteins to compensate for the sequence information. Particularly, a S-C cross-modal cross-attention fusion network module is proposed that not only optimizes the weights of the semantic information but also efficiently fuses the sequence features with the structural information, thus avoiding the simple splicing of different modal features. Our experimental results demonstrate that the proposed DMMAFS outperforms the state-of-the-art methods in protein function prediction. Liangwen He, Zhaohong Deng, Fuping Hu, Yun Zuo 0001, Haoran Chen 0003, Xiaoyong Pan, Zhisheng Wei, Hong-Bin Shen, Dongjun Yu, Jing Wu 0030 |
IEEE Trans. Comput. Biol. Bioinform. | 2 |
| 2025 | SEFP: Structure-Based Enzyme Function PredictionabstractTraditional biological experimental methods to determine enzyme properties are time-consuming and costly, leading to an increasing interest in computational models for enzyme function prediction. However, the existing computational methods are insufficient and inefficient to exploit enzyme structure. In this work, we introduce SEFP, a novel method leveraging enzyme point clouds for enzyme function prediction. The structure encoder of SEFP uses a tailored enzyme point cloud network to analyze the three-dimensional arrangement of atoms within the enzyme, integrating hierarchical residue global features through a residue feature adapter to extract detailed enzyme point features. Additionally, the Bio-BCS residue feature encoder extracts enzyme residue features with channel and spatial weights using a specially designed attention mechanism. Finally, SEFP fuses point and residue features to generate the final prediction results. Comparative evaluations show that SEFP outperforms various recent computational methods, demonstrating superior performance. On the RSCB enzyme structure dataset, SEFP achieves an f1-score of 95.85, outperforming two representative structure-based methods, EnzyNet and DeepFri. On the HECNet dataset, SEFP maintains its superiority over all comparison sequence-based methods, yielding an f1-score of 94.29. Ablation studies are conducted to confirm the effectiveness of individual modules within SEFP. These findings underscore the potential of SEFP for reliable and precise enzyme function prediction, offering advancements in bioinformatics and computational biology. Guanqing Yu, Zhaohong Deng, Chenxi Luo, Cheng Cai, Wei Zhang 0221, Fuping Hu, Kup-Sze Choi, Zhisheng Wei, Jing Wu 0030 |
IEEE Trans. Comput. Biol. Bioinform. | 3 |
| 2025 | Automated Cluster Elimination Guided by High-Density PointsabstractDetermining the optimal number of clusters in cluster analysis without prior knowledge remains a critical and challenging task. Existing methods often depend on calculating clustering validity indices (CVIs), which increases complexity and may reduce efficiency. Furthermore, different CVIs frequently suggest varying optimal cluster numbers, complicating the selection process. To address these challenges, we propose a novel clustering algorithm, self-regulating possibilistic C-means (PCM) with high-density points (SR-PCM-HDP), which simplifies cluster number determination while improving clustering efficiency. First, the density-based knowledge extraction (DBKE) method is introduced to estimate an appropriate initial cluster number and identify high-density points. DBKE enhances the density peak clustering (DPC) algorithm by removing the need for a predefined density radius. Second, SR-PCM-HDP refines the clustering process by incorporating a parameter to balance the interactions between high-density points and cluster centers, reducing sensitivity to initial configurations and accelerating convergence. Third, the parameter adjustment mechanism in classical PCM is redefined to enable adaptive updates during SR-PCM-HDP iterations. This mechanism facilitates the gradual elimination of obsolete clusters and iterative cluster formation. The theoretical foundations of the SR-PCM-HDP cluster elimination mechanism are rigorously established. Experimental results validate the accuracy and effectiveness of SR-PCM-HDP in determining cluster numbers and ensuring clustering validity, particularly for datasets with overlapping or imbalanced distributions. Comparisons are conducted against 13 state-of-the-art algorithms, including fuzzy clustering, possibilistic clustering, and CVI-based cluster determination methods. Xianghui Hu, Yichuan Jiang, Witold Pedrycz, Zhaohong Deng, Jianwei Gao, Yiming Tang 0001 |
IEEE Trans. Cybern. | 4 |
| 2025 | One-Step Fuzzy Ensemble Clustering Method via Embedding Ground-Truth Cluster Number GraphsabstractIn multisource clustering tasks, the number of clusters in each source or view may not align with the number of ground-truth clusters. Existing ensemble clustering methods face two notable challenges: (1) developing a new ensemble framework that yields a final clustering result matching the ground-truth cluster count and (2) revealing consistency among all base clustering results. To address these challenges, we propose a novel one-step fuzzy ensemble clustering method (OS-FECM) that incorporates ground-truth cluster number graphs. Initially, OS-FECM establishes a one-step fuzzy ensemble framework that directly integrates all base fuzzy clustering results (i.e., membership matrices) with varying cluster counts, thereby eliminating reliance on the CA matrix typical of existing two-step ensemble frameworks. Furthermore, we construct a ground-truth cluster number graph, which maps the number of clusters in each base clustering result to the ground-truth cluster count in the final ensemble result. This graph reveals the consistency among all base fuzzy clustering results and illustrates the relationships between clusters in the base results and the ground-truth clusters. It is then embedded into the corresponding base fuzzy clustering results to enhance the final ensemble result. Lastly, we employ an alternating optimization method alongside a weighting mechanism to derive the final ensemble clustering result and adaptively assign importance to each base clustering result. Experimental evaluations across various datasets demonstrate that OS-FECM achieves clustering performance that is at least comparable to, if not superior to, that of other comparative methods. Zekang Bian, Zhaohong Deng, Shitong Wang 0001 |
IEEE Trans. Fuzzy Syst. | 3 |
| 2025 | FHN: Fuzzy Hashing Network for Medical Image RetrievalabstractThe rapid advancement of medical imaging technologies has led to an exponential increase in medical image data, making efficient retrieval from large-scale datasets critical for improving diagnostic accuracy and speed. However, two key challenges hinder this process: first, the presence of uncertain and subtle lesions in medical images that are often difficult to discern, and second, class imbalance across different case types within medical image databases. These inherent challenges significantly degrade the performance of existing hashing algorithms. In recent years, methods based on the Takagi–Sugeno–Kang fuzzy system (TSK-FS) have shown promising performance in medical image modeling. Inspired by these advances, this article proposes a novel fuzzy hashing network (FHN) based on TSK-FS to enhance retrieval performance by effectively handling both uncertainty and data imbalance in medical imaging. The FHN first introduces a novel fuzzification mechanism that incorporates the concept of a self-attention mechanism to effectively capture the complex underlying features in medical images, thereby enhancing the data discriminability in fuzzy spaces. Meanwhile, a new consequent parameter learning mechanism is developed for defuzzification by introducing the Transformer network, which aims to improve the inference efficiency and generalization capability of the FHN. Based on these two mechanisms, FHN's capability of analyzing and handling uncertain data is significantly enhanced. Furthermore, a novel hash center loss is designed to capture global relationships while emphasizing local structural information, thereby improving the handling of imbalanced data and significantly enhancing retrieval performance. Weiping Ding 0001, Linlin Zhou, Wei Zhang 0221, Te Zhang, Zhaohong Deng, Yuanpeng Zhang 0001, Guanjin Wang |
IEEE Trans. Fuzzy Syst. | 5 |
| 2025 | Robust Federated Fuzzy C-Means Algorithm in Heterogeneous ScenariosabstractThe federated Fuzzy C-means (federated FCM) extends the traditional Fuzzy C-means (FCM) to the federated learning (FL) scenario, aiming to address the data privacy preservation issue of soft clustering in distributed environments. However, a significant challenge persists with existing federated FCM algorithms, i.e., they struggle to converge effectively in complex heterogeneous scenarios, leading to unstable clustering outcomes. Here the complex heterogeneous scenarios stem from the combination of non-independently and identically distributed (non-IID) data across different clients (statistical heterogeneity), coupled with the involvement of only some clients in each iteration (systematic heterogeneity). While prior research has attempted to address the impact of statistical heterogeneity in FL scenarios, it has overlooked the issue of system heterogeneity. In response, this paper proposes a novel federated FCM algorithm (SC-FFCM) that remains robust even in such complex heterogeneous scenarios. Firstly, the client-side clustering module of SC-FFCM adopts a Gradient-Based FCM algorithm, facilitating corrections to the direction of local optimization. Secondly, the algorithm introduces a control variates technique to rectify update bias during the iteration process, thereby mitigating the adverse effects of random client sampling and non-IID data distribution on the algorithm convergence. Finally, the proposed algorithm approximates the ideal federated FCM algorithm. Experimental studies verify the effectiveness of the proposed method. The source code of the proposed SC-FFCM algorithm is available from the following website https://github.com/Creazy-MR/SC-FFCM. Qixian Zhang, Zhaohong Deng, Wei Zhang 0221, Zhuangzhuang Zhao, Zhiyong Xiao 0001, Kup-Sze Choi, Guanjin Wang, Yuxi Ge, Shudong Hu |
IEEE Trans. Fuzzy Syst. | 2 |
| 2025 | Dual Anchor Graph Fuzzy Clustering for Multiview DataabstractMultiview anchor graph clustering has been a prominent research area in recent years, leading to the development of several effective and efficient methods. However, three challenges are faced by current multiview anchor graph clustering methods. First, real-world data often exhibit uncertainty and poor discriminability, leading to suboptimal anchor graphs when directly extracted from the original data. Second, most existing methods assume the presence of common information between views and primarily explore it for clustering, thus neglecting view-specific information. Third, further exploration and exploitation of the learned anchor graph to enhance clustering performance remains an open research question. To address these issues, a novel dual anchor graph fuzzy clustering method is proposed in this article. First, a novel matrix factorization-based dual anchor graph learning method is proposed to address the first two issues by extracting highly discriminative hidden representations for each view and subsequently deriving both common and specific anchor graphs from these hidden representations. Then, to address the third issue, a novel anchor graph fuzzy clustering method is developed with cooperative learning to exploit and utilize the common and specific anchor graphs fully. Meanwhile, a fuzzy membership structure preservation mechanism with dual anchor graphs is constructed to enhance clustering performance. Finally, negative Shannon entropy is further introduced to adaptively adjust the view weighing. Extensive experiments on several datasets demonstrate the effectiveness of the proposed method. Wei Zhang 0221, Xiuyu Huang, Andong Li, Te Zhang, Weiping Ding 0001, Zhaohong Deng, Shitong Wang 0001 |
IEEE Trans. Fuzzy Syst. | 6 |
| 2025 | From Local to Global: A Progressive Collaborative Learning Framework for Multitask TSK Fuzzy System ModelingabstractMulti-task Takagi-Sugeno-Kang fuzzy systems (MT-TSK-FS) commonly utilize interpretable fuzzy rules to facilitate information sharing among tasks and have shown promising performance. However, existing MT-TSK-FS modeling approaches still face several challenges. Specifically, they mainly focus on the global fitting of tasks while overlooking the local fitting accuracy of individual rules, thereby undermining rule interpretability. Moreover, the antecedent and consequent components of fuzzy rules are often learned independently, lacking collaborative coordination. In addition, task-specific rule diversity is insufficiently addressed, which may result in redundant or overlapping rules. To this end, we propose a Progressive Collaborative Learning framework for MT-TSK-FS (MTTSKFS-PCL), which progressively optimizes the model from the local perspective of individual rules to that of global tasks. At the local level, a Locally Target-Guided Multi-Task Fuzzy Clustering Method is introduced to enhance local rule accuracy and to jointly optimize rule antecedents and consequents. At the global level, a Mini-Batch Gradient Descent (MBGD) algorithm is employed to integrate fuzzy rules and enhance overall task fitting. Meanwhile, a confidence-supervised regularization strategy is introduced to preserve local rule accuracy throughout the MBGD process. Furthermore, a rule diversity enhancement mechanism is incorporated to improve task-specific rule diversity. Extensive experimental evaluations on multiple benchmark datasets demonstrate that the proposed MTTSKFS-PCL significantly outperforms existing state-of-the-art methods, validating its effectiveness and robustness in multi-task fuzzy modeling. Zhuangzhuang Zhao, Zhaohong Deng, Chenxi Luo, Kup-Sze Choi, Shitong Wang 0001, Yuxi Ge, Shudong Hu |
IEEE Trans. Fuzzy Syst. | 2 |
| 2025 | No-Reference Image Quality Assessment Leveraging GenAI ImagesabstractIn recent years, deep learning-based methods have made significant progress on the image quality assessment problem; however, challenges remain arising from the lack of annotated, real-world training data and consequent poor generalization ability. Towards addressing these challenges, we propose a no-reference image quality assessment (NR-IQA) method based on generative AI (GenAI) images. Specifically, we use GenAI images as reference images, employing a cold diffusion model to generate distorted images of four different distortion types, and we label these distorted images using a full-reference model, thereby making it possible to construct a large-scale pre-training dataset. We use this resource generation method to facilitate NR-IQA model building. We deploy a Multi-scale Cross Attention Block (MCAB) and a Scale Simple Attention Module (SSAM) to enhance feature representation by extracting multi-scale feature information from both the channel and spatial dimensions that are predictive of image quality. Extensive experiments on eight public databases demonstrate that the proposed method achieves state-of-the-art (SOTA) performance. A public release of all the codes associated with this work will be made available on GitHub. Qingbing Sang, Qian Li 0060, Lixiong Liu, Zhaohong Deng, Xiaojun Wu 0001, Alan C. Bovik |
IEEE Trans. Image Process. | 4 |
| 2025 | MBSCLoc: Multi-Label Subcellular Localization Predict Based on Cluster Balanced Subspace Partitioning Method and Multi-Class Contrastive Representation LearningabstractmRNA subcellular localization is a prevalent and essential mechanism that precisely regulates protein translation and significantly impacts various cellular processes. mRNA subcellular localization has advanced the understanding of mRNA function, yet existing methods face limitations, including imbalanced data, suboptimal model performance, and inadequate generalization, particularly in multi-label localization scenarios where solutions are scarce. This study introduces MBSCLoc, a predictor for mRNA multi-label subcellular localization. MBSCLoc predicts mRNA locations across multiple cellular compartments simultaneously, overcoming challenges like single-location prediction, incomplete feature extraction, and imbalanced data. MBSCLoc leverages UTR-LM model for feature extraction, followed by multi-class contrastive representation learning and Clustering Balanced Subspace Partitioning to construct balanced subspaces. It then optimizes sample distribution to tackle severe data imbalance and uses multiple XGBoost classifiers, integrated through voting, to enhance accuracy and generalization. Five-fold cross-validation and independent testing results show that MBSCLoc significantly outperforms other methods. Additionally, MBSCLoc offers superior pixel-level interpretability, strongly supporting mRNA multi-label subcellular localization research. Crucially, the importance of the 5' UTR and 3' UTR regions has been preliminarily confirmed using traditional biological analysis and Tree-SHAP, with most mRNA sequences showing significant relevance in these regions, especially the 3' UTR where about 80% of specific sites reach peak significance. Bangyi Zhang, Yun Zuo 0001, Zhiqiang Dai, Sifan Zhu, Zhaohong Deng |
IEEE J. Biomed. Health Informatics | 6 |
| 2025 | LM-TLIPs: Integrating Large Model and Transfer Learning Technology for Precise Identification of Phosphorylation Sites in SARS-CoV-2abstractIn recent years, the rapid spread of SARS-CoV-2 has triggered a global health crisis and socio-economic challenges. As a crucial post-translational modification, phosphorylation plays a vital role in the regulation of cellular functions. Given its close relationship with SARS-CoV-2 infection, accurately identifying virus-induce phosphorylation sites is essential for understanding the molecular mechanisms of viral infection and its impact on host cells. Although the development of various computational tools for predicting phosphorylation sites, these tools have several shortcomings, such as insufficient data and limited model generalization ability, which limit their effectiveness in practical applications. To overcome these limitations, this study proposes a novel method for predicting SARS-CoV-2 phosphorylation sites, LM-TLIPs, based on the latest technology. This method uses the most advanced large model technology ESM-2 to extract information from S/T sites and Y sites; by fine-tuning the large model and introducing transfer learning technology, it addresses the challenge of accurately predicting Y sites due to insufficient data in this study. Independent testing on S/T sites(Acc:0.8309, Sn:0.8443, Sp:0.8174, MCC:0.6620, AUC:0.8993) and Y sites(Acc:0.9048, Sn:0.9524, Sp:0.8571, MCC:0.8132, AUC:0.9388) has validated that LM-TLIPs outperforms existing optimal prediction tools, demonstrating its superior ability in identifying phosphorylation sites. Furthermore, we conducted an exhaustive interpretability analysis based on attention weight heatmaps and feature importance ranking to enhance the transparency and confidence of prediction results. Yun Zuo 0001, Minquan Wan, Xinyue Shao, Dandan Qiao, Bulanni Xiong, Zhaohong Deng |
IEEE J. Biomed. Health Informatics | 6 |
| 2025 | Clustering Interval and Triangular Granular Data: Modeling, Execution, and AssessmentabstractIn current granular clustering algorithms, numeric representatives were selected by users or an ordinary strategy, which seemed simple; meanwhile, weight settings for granular data could not adequately express their structural characteristics. Aiming at these problems, in this study, a new scheme called a granular weighted kernel fuzzy clustering (GWKFC) algorithm is put forward. We propose the representative selection and granularity generation (RSGG) algorithm enlightened by the density peak clustering (DPC) algorithm. We build interval and triangular granular data on the strength of numeric representatives obtained by RSGG under the principle of justifiable granularity (PJG), in which we establish some combinations of functions and boundary constraints and prove their properties. Furthermore, we present a novel distance formula via the kernel function for granular data and design new weights to affect the coverage and specificity of granular data. In addition, based upon these factors, we come up with the GWKFC algorithm of granular clustering, and its performance with different granularity is assessed. To sum up, a macro framework involving granular modeling, granular clustering, and assessment has been set up. Lastly, the GWKFC algorithm and ten other granular clustering algorithms are compared by experiments on some artificial and UCI datasets together with datasets with large data or those of high dimensionality. It is found that the GWKFC algorithm can provide better granular clustering results by contrast with other algorithms. The originality is embodied as follows. First, we improve the previous density radius and present the RSGG algorithm to acquire numeric representatives. Second, we propose a new strategy to determine granular data boundaries and further obtain novel weights enlightened by the idea of volume. Lastly, we employ the kernel function to calculate the distance between granular data, which has a stronger spatial division ability than the previous Euclidean distance. Yiming Tang 0001, Witold Pedrycz, Jianwei Gao, Xianghui Hu, Zhaohong Deng |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2024 | MSlocPRED: deep transfer learning-based identification of multi-label mRNA subcellular localizationabstractSubcellular localization of messenger ribonucleic acid (mRNA) is a universal mechanism for precise and efficient control of the translation process. Although many computational methods have been constructed by researchers for predicting mRNA subcellular localization, very few of these computational methods have been designed to predict subcellular localization with multiple localization annotations, and their generalization performance could be improved. In this study, the prediction model MSlocPRED was constructed to identify multi-label mRNA subcellular localization. First, the preprocessed Dataset 1 and Dataset 2 are transformed into the form of images. The proposed MDNDO-SMDU resampling technique is then used to balance the number of samples in each category in the training dataset. Finally, deep transfer learning was used to construct the predictive model MSlocPRED to identify subcellular localization for 16 classes (Dataset 1) and 18 classes (Dataset 2). The results of comparative tests of different resampling techniques show that the resampling technique proposed in this study is more effective in preprocessing for subcellular localization. The prediction results of the datasets constructed by intercepting different NC end (Both the 5' and 3' untranslated regions that flank the protein-coding sequence and influence mRNA function without encoding proteins themselves.) lengths show that for Dataset 1 and Dataset 2, the prediction performance is best when the NC end is intercepted by 35 nucleotides, respectively. The results of both independent testing and five-fold cross-validation comparisons with established prediction tools show that MSlocPRED is significantly better than established tools for identifying multi-label mRNA subcellular localization. Additionally, to understand how the MSlocPRED model works during the prediction process, SHapley Additive exPlanations was used to explain it. The predictive model and associated datasets are available on the following github: https://github.com/ZBYnb1/MSlocPRED/tree/main. Yun Zuo 0001, Bangyi Zhang, Wenying He, Yue Bi, Xiangrong Liu, Xiangxiang Zeng, Zhaohong Deng |
Briefings Bioinform. | 7 |
| 2024 | MINDG: a drug-target interaction prediction method based on an integrated learning algorithmabstractMOTIVATION: Drug-target interaction (DTI) prediction refers to the prediction of whether a given drug molecule will bind to a specific target and thus exert a targeted therapeutic effect. Although intelligent computational approaches for drug target prediction have received much attention and made many advances, they are still a challenging task that requires further research. The main challenges are manifested as follows: (i) most graph neural network-based methods only consider the information of the first-order neighboring nodes (drug and target) in the graph, without learning deeper and richer structural features from the higher-order neighboring nodes. (ii) Existing methods do not consider both the sequence and structural features of drugs and targets, and each method is independent of each other, and cannot combine the advantages of sequence and structural features to improve the interactive learning effect. RESULTS: To address the above challenges, a Multi-view Integrated learning Network that integrates Deep learning and Graph Learning (MINDG) is proposed in this study, which consists of the following parts: (i) a mixed deep network is used to extract sequence features of drugs and targets, (ii) a higher-order graph attention convolutional network is proposed to better extract and capture structural features, and (iii) a multi-view adaptive integrated decision module is used to improve and complement the initial prediction results of the above two networks to enhance the prediction performance. We evaluate MINDG on two dataset and show it improved DTI prediction performance compared to state-of-the-art baselines. AVAILABILITY AND IMPLEMENTATION: https://github.com/jnuaipr/MINDG. Hailong Yang 0001, Yun Zuo 0001, Zhaohong Deng, Xiaoyong Pan, Hong-Bin Shen, Kup-Sze Choi, Dongjun Yu |
Bioinform. | 4 |
| 2024 | PreMLS: The undersampling technique based on ClusterCentroids to predict multiple lysine sitesabstractThe translated protein undergoes a specific modification process, which involves the formation of covalent bonds on lysine residues and the attachment of small chemical moieties. The protein's fundamental physicochemical properties undergo a significant alteration. The change significantly alters the proteins' 3D structure and activity, enabling them to modulate key physiological processes. The modulation encompasses inhibiting cancer cell growth, delaying ovarian aging, regulating metabolic diseases, and ameliorating depression. Consequently, the identification and comprehension of post-translational lysine modifications hold substantial value in the realms of biological research and drug development. Post-translational modifications (PTMs) at lysine (K) sites are among the most common protein modifications. However, research on K-PTMs has been largely centered on identifying individual modification types, with a relative scarcity of balanced data analysis techniques. In this study, a classification system is developed for the prediction of concurrent multiple modifications at a single lysine residue. Initially, a well-established multi-label position-specific triad amino acid propensity algorithm is utilized for feature encoding. Subsequently, PreMLS: a novel ClusterCentroids undersampling algorithm based on MiniBatchKmeans was introduced to eliminate redundant or similar major class samples, thereby mitigating the issue of class imbalance. A convolutional neural network architecture was specifically constructed for the analysis of biological sequences to predict multiple lysine modification sites. The model, evaluated through five-fold cross-validation and independent testing, was found to significantly outperform existing models such as iMul-kSite and predML-Site. The results presented here aid in prioritizing potential lysine modification sites, facilitating subsequent biological assays and advancing pharmaceutical research. To enhance accessibility, an open-access predictive script has been crafted for the multi-label predictive model developed in this study. Yun Zuo 0001, Xingze Fang, Jiayong Wan, Wenying He, Xiangrong Liu, Xiangxiang Zeng, Zhaohong Deng |
PLoS Comput. Biol. | 7 |
| 2024 | MVDINET: A Novel Multi-Level Enzyme Function Predictor With Multi-View Deep Interactive LearningabstractAs a class of extremely significant of biocatalysts, enzymes play an important role in the process of biological reproduction and metabolism. Therefore, the prediction of enzyme function is of great significance in biomedicine fields. Recently, computational methods for predicting enzyme function have been proposed, and they effectively reduce the cost of enzyme function prediction. However, there are still deficiencies for effectively mining the discriminant information for enzyme function recognition in existing methods. In this study, we present MVDINET, a novel method for multi-level enzyme function prediction. First, the initial multi-view feature data is extracted by the enzyme sequence. Then, the above initial views are fed into various deep specific network modules to learn the depth-specificity information. Further, a deep view interaction network is designed to extract the interaction information. Finally, the specificity information and interaction information are fed into a multi-view adaptively weighted classification. We compressively evaluate MVDINET on benchmark datasets and demonstrate that MVDINET is superior to existing methods. Wenliang Tang, Zhaohong Deng, Hanwen Zhou, Wei Zhang 0221, Fuping Hu, Kup-Sze Choi, Shitong Wang 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2024 | HGLA: Biomolecular Interaction Prediction Based on Mixed High-Order Graph Convolution With Filter Network via LSTM and Channel AttentionabstractPredicting biomolecular interactions is significant for understanding biological systems. Most existing methods for link prediction are based on graph convolution. Although graph convolution methods are advantageous in extracting structure information of biomolecular interactions, two key challenges still remain. One is how to consider both the immediate and high-order neighbors. Another is how to reduce noise when aggregating high-order neighbors. To address these challenges, we propose a novel method, called mixed high-order graph convolution with filter network via LSTM and channel attention (HGLA), to predict biomolecular interactions. Firstly, the basic and high-order features are extracted respectively through the traditional graph convolutional network (GCN) and the two-layer Higher-Order Graph Convolutional Architectures via Sparsified Neighborhood Mixing (MixHop). Secondly, these features are mixed and input into the filter network composed of LayerNorm, SENet and LSTM to generate filtered features, which are concatenated and used for link prediction. The advantages of HGLA are: 1) HGLA processes high-order features separately, rather than simply concatenating them; 2) HGLA better balances the basic features and high-order features; 3) HGLA effectively filters the noise from high-order neighbors. It outperforms state-of-the-art networks on four benchmark datasets. Zhaohong Deng, Ruibo Li, Wei Zhang 0221, Qiongdan Lou, Kup-Sze Choi, Shitong Wang 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2024 | Graph Fuzzy System for the Whole Graph Prediction: Concepts, Models, and AlgorithmsabstractFuzzy systems (FSs) have been widely utilized in diverse domains, such as pattern recognition, intelligent control, data mining, and bioinformatics due to their strong interpretation and learning abilities. Traditionally, FSs have mainly been applied to model Euclidean data. However, with the emergence of scenarios involving graph data, such as social networks and traffic route maps, which inherently possess non-Euclidean structures, there is a need to develop FS modeling methods suitable for graph data while retaining the advantages of traditional FSs. This article presents a novel FS called graph fuzzy system (GFS) specifically designed for modeling whole graph data. The concepts, modeling framework, and construction algorithms are systematically developed. First, the article defines GFS-related concepts, including the graph fuzzy rule base, graph fuzzy sets, and graph consequent processing unit (GCPU). Second, the learning framework for GFS is proposed. It includes a novel K-Means with graph similarity measure clustering approach (KM-GSM) for generating antecedents in GFS and a new consequent parameters learning algorithm based on graph neural network (GNN). Moreover, three different versions of the GFS implementation algorithm are developed and thoroughly evaluated through experiments on various graph prediction datasets. The results demonstrate that the proposed GFS inherits the advantages of mainstream GNNs methods and conventional FSs methods while achieving superior performance in whole graph prediction compared to existing approaches. Fuping Hu, Zhaohong Deng, Guanjin Wang, Zhenping Xie, Kup-Sze Choi, Shitong Wang 0001 |
IEEE Trans. Fuzzy Syst. | 2 |
| 2024 | GFS-Node: Graph Fuzzy Systems for Node PredictionabstractGraph data modeling is nontrivial due to the challenges to ensure model interpretability and handle data uncertainty. While methods derived from deep learning models, such as graph neural networks (GNNs), are able to handle graph data, the interpretability is limited. Graph fuzzy systems (GFSs) based on the fuzzy rules and fuzzy inference have been proposed to improve interpretability, but the existing methods are developed for whole graph prediction only and cannot deal with node prediction, which is a more common task in graph data modeling. To tackle the challenges, a novel GFS for node prediction (GFS-node) is investigated in this study. For this purpose, the concepts, framework, and algorithms of GFS-node are systematically developed. First, several related concepts are defined, including the node fuzzy rule base, node fuzzy set, and node consequent processing module (NCPM). A general framework for GFS-node is then presented, where the construction of antecedents and the consequents of fuzzy rules are analyzed. Furthermore, a concrete implementation method of GFS-node is designed. In particular, the kernelKvirtual central nodes clustering (KVCN) algorithm is proposed to develop the algorithm for antecedent generation, and the linear message passing network (LMPN) is adopted to develop the algorithm for consequent generation and learning. Experiments are carried out on multiple benchmark datasets, and the results show that GFS-node combines the advantages of both traditional fuzzy systems and classical GNNs for node prediction. Fuping Hu, Zhaohong Deng, Zhenping Xie, Te Zhang, Kup-Sze Choi, Shitong Wang 0001 |
IEEE Trans. Fuzzy Syst. | 2 |
| 2024 | Multiview Transfer Representation Learning With TSK Fuzzy System for EEG Epilepsy DetectionabstractAutomatic analysis of epileptic encephalography signals with intelligent models can greatly reduce the workload of doctors. However, the lack of data, insufficient labels, and inconsistent data distribution in real-world scenarios significantly affect the performance of intelligent models. Transfer learning plays an important role in solving the above problems but some challenges remain. First, while various feature extraction methods are available to extract features from the original epilepsy signal, it is difficult to determine which features are effective. Second, transfer learning may lead to domain information loss since the original feature representation from different domains is changed. Third, most of the existing models lack transparency to provide medical practitioners confidence of use. To this end, this article proposes the novel method Multiview Information Preservation Transfer Representation Learning based on Fuzzy Systems (MIP-TRL-FS) to address the issues. First, MIP-TRL-FS utilizes multiple views to get rid of the feature selection process. Second, information preservation techniques are utilized to maintain the data information from the aspects of sample level and feature level, thus minimizing information loss during the transfer learning process. Third, by using Takagi–Sugeno–Kang fuzzy systems as the base model, the output of the proposed method can be interpreted linguistically with IF-THEN rules to makes the model transparent. Extensive experiments were conducted on the CHB-MIT dataset and the results demonstrate the effectiveness of the proposed method. Andong Li, Zhaohong Deng, Wei Zhang 0221, Zhiyong Xiao 0001, Kup-Sze Choi, Shudong Hu, Shitong Wang 0001 |
IEEE Trans. Fuzzy Syst. | 2 |
| 2024 | Rules-Based Heterogeneous Feature Transfer Learning Using Fuzzy InferenceabstractHeterogeneous feature transfer (HeFT) learning can leverage the semantically related source domain from a different feature space for modeling the target domain with insufficient information. Although HeFT learning has made significant progress, it still faces two major challenges: weak interpretability of the transfer process and underutilization of the hidden information of the heterogeneous source and target domains. To address these two challenges, a framework called heterogeneous feature transfer using fuzzy inference rules (HeFT-FIR) is proposed. The HeFT-FIR framework has two parts: First, design of Takagi–Sugeno–Kang fuzzy systems (TSK-FSs) for the source and target domains, respectively, to achieve HeFT and enhance the interpretability of the transfer process; and second, integration of the HeFT learning mechanism with fuzzy inference rules to optimize the parameters of TSK-FSs and mine the hidden information of the two domains. Based on the framework, a TSK-FS-based heterogeneous feature transfer learning method is then developed with three fuzzy feature space-based learning mechanisms for joint distribution adaptation, local geometric property preservation, and heterogeneous discriminant information extraction, respectively. The mechanisms reduce the difference in distribution between the heterogeneous source and target domains in a common feature subspace, preserve the local geometric properties of two domains, and extract the global discriminant information of them. Extensive analyses are conducted to verify the superiority of the proposed framework and method. Qiongdan Lou, Wu Sun, Wei Zhang 0221, Zhaohong Deng, Kup-Sze Choi, Shitong Wang 0001 |
IEEE Trans. Fuzzy Syst. | 4 |
| 2024 | End-to-End Multiview Fuzzy Clustering With Double Representation Learning and Visible-Hidden View CooperationabstractMultiview clustering has received great attention in recent years for the potential in clustering performance improvement by using cooperative learning of different views. Despite the considerable progress, a few issues remain: 1) real multiview data contains redundant features and noises that lead to unsatisfactory clustering performance; 2) most existing multiview clustering methods only mine the shared information between views and ignore the specific information within views; and 3) most multiview clustering methods are based on a two-step framework that learn the hidden view representation and then perform clustering, overlooking the correlation between the two processes. Although some approaches have been proposed to deal with these issues, they cannot them simultaneously. To this end, we propose an end-to-end multiview fuzzy clustering. First, we construct a multiview fuzzy clustering framework to mine the specific information of the visible views. Second, to reduce the impact of redundant features and noises on clustering performance, we introduce the orthogonal projection matrix into the clustering framework to learn the low-dimensional representation of the visible views. Meanwhile, this procedure is integrated into the clustering framework. Third, we explore the shared hidden view representation between the visible views by multiview non-negative matrix factorization and integrate it into the clustering framework to realize visible-hidden view cooperation learning. Finally, the shared hidden view representation learning between visible views, the low-dimensional representation learning of visible views, and the clustering partition of multiview data negotiate with each other in the end-to-end learning framework. Extensive experiments on benchmark multiview datasets indicate the superiority of the proposed method over state-of-the-art methods. Hongtan Yang, Zhaohong Deng, Wei Zhang 0221, Qunzhuo Wu, Kup-Sze Choi, Shitong Wang 0001 |
IEEE Trans. Fuzzy Syst. | 2 |
| 2024 | Pseudolabel Enhanced Multiview Deep Concept Factorization Fuzzy ClusteringabstractMultiview fuzzy c-means clustering has garnered significant attention in recent years, leading to the development of various multiview fuzzy clustering algorithms. However, existing algorithms still exhibit room for improvement. First, most existing algorithms only utilize the shallow information of the view data and fail to delve into the mining and utilization of deeper representations. Second, existing algorithms tend to extract common representations among the views first and then implement clustering separately, which may lack a collaborative linkage between two tasks. Finally, multiview clustering algorithms based on representation learning often overlook the importance of effectively preserving similarity information within the views. To address these limitations, we propose a novel algorithm called pseudolabel enhanced multiview deep concept factorization fuzzy clustering (PE-MV-DCFCM). The algorithm first introduces a deep concept factorization method to uncover the deep information of the view data. Subsequently, it employs pseudolabel learning to preserve intraview similarity information during the learning of common representations among the views, based on non-negative matrix factorization. Finally, this algorithm integrates deep concept factorization, representation learning, and fuzzy clustering into a unified framework to enhance the collaboration among the various substeps of the algorithm. Experiments on several benchmark datasets show that the proposed PE-MV-DCFCM algorithm outperformed other state-of-the-art algorithms. Zhuangzhuang Zhao, Hongtan Yang, Zhaohong Deng, Wei Zhang 0221, Chenxi Luo, Guanjin Wang, Yuxi Ge, Shudong Hu |
IEEE Trans. Fuzzy Syst. | 3 |
| 2024 | Multi-View Fuzzy Representation Learning With Rules Based ModelabstractUnsupervised multi-view representation learning has been extensively studied for mining multi-view data. However, some critical challenges remain. On the one hand, the existing methods cannot explore multi-view data comprehensively since they usually learn a common representation between views, given that multi-view data contains both the common information between views and the specific information within each view. On the other hand, to mine the nonlinear relationship between data, kernel or neural network methods are commonly used for multi-view representation learning. However, these methods are lacking in interpretability. To this end, this paper proposes a new multi-view fuzzy representation learning method based on the interpretable Takagi-Sugeno-Kang (TSK) fuzzy system (MVRL_FS). The method realizes multi-view representation learning from two aspects. First, multi-view data are transformed into a high-dimensional fuzzy feature space, while the common information between views and specific information of each view are explored simultaneously. Second, a new regularization method based on L2,1-norm regression is proposed to mine the consistency information between views, while the geometric structure of the data is preserved through the Laplacian graph. Finally, extensive experiments on many benchmark multi-view datasets are conducted to validate the superiority of the proposed method. Wei Zhang 0221, Zhaohong Deng, Te Zhang, Kup-Sze Choi, Shitong Wang 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2024 | One-Step Multiview Fuzzy Clustering With Collaborative Learning Between Common and Specific Hidden Space InformationabstractMultiview data are widespread in real-world applications, and multiview clustering is a commonly used technique to effectively mine the data. Most of the existing algorithms perform multiview clustering by mining the commonly hidden space between views. Although this strategy is effective, there are two challenges that still need to be addressed to further improve the performance. First, how to design an efficient hidden space learning method so that the learned hidden spaces contain both shared and specific information of multiview data. Second, how to design an efficient mechanism to make the learned hidden space more suitable for the clustering task. In this study, a novel one-step multiview fuzzy clustering (OMFC-CS) method is proposed to address the two challenges by collaborative learning between the common and specific space information. To tackle the first challenge, we propose a mechanism to extract the common and specific information simultaneously based on matrix factorization. For the second challenge, we design a one-step learning framework to integrate the learning of common and specific spaces and the learning of fuzzy partitions. The integration is achieved in the framework by performing the two learning processes alternately and thereby yielding mutual benefit. Furthermore, the Shannon entropy strategy is introduced to obtain the optimal views weight assignment during clustering. The experimental results based on benchmark multiview datasets demonstrate that the proposed OMFC-CS outperforms many existing methods. Wei Zhang 0221, Zhaohong Deng, Te Zhang, Kup-Sze Choi, Shitong Wang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | MMSMAPlus: a multi-view multi-scale multi-attention embedding model for protein function predictionabstractProtein is the most important component in organisms and plays an indispensable role in life activities. In recent years, a large number of intelligent methods have been proposed to predict protein function. These methods obtain different types of protein information, including sequence, structure and interaction network. Among them, protein sequences have gained significant attention where methods are investigated to extract the information from different views of features. However, how to fully exploit the views for effective protein sequence analysis remains a challenge. In this regard, we propose a multi-view, multi-scale and multi-attention deep neural model (MMSMA) for protein function prediction. First, MMSMA extracts multi-view features from protein sequences, including one-hot encoding features, evolutionary information features, deep semantic features and overlapping property features based on physiochemistry. Second, a specific multi-scale multi-attention deep network model (MSMA) is built for each view to realize the deep feature learning and preliminary classification. In MSMA, both multi-scale local patterns and long-range dependence from protein sequences can be captured. Third, a multi-view adaptive decision mechanism is developed to make a comprehensive decision based on the classification results of all the views. To further improve the prediction performance, an extended version of MMSMA, MMSMAPlus, is proposed to integrate homology-based protein prediction under the framework of multi-view deep neural model. Experimental results show that the MMSMAPlus has promising performance and is significantly superior to the state-of-the-art methods. The source code can be found at https://github.com/wzy-2020/MMSMAPlus. Zhaohong Deng, Wei Zhang 0221, Qiongdan Lou, Kup-Sze Choi, Zhisheng Wei, Jing Wu 0030 |
Briefings Bioinform. | 2 |
| 2023 | MLNGCF: circRNA-disease associations prediction with multilayer attention neural graph-based collaborative filteringabstractMOTIVATION: CircRNAs play a critical regulatory role in physiological processes, and the abnormal expression of circRNAs can mediate the processes of diseases. Therefore, exploring circRNAs-disease associations is gradually becoming an important area of research. Due to the high cost of validating circRNA-disease associations using traditional wet-lab experiments, novel computational methods based on machine learning are gaining more and more attention in this field. However, current computational methods suffer to insufficient consideration of latent features in circRNA-disease interactions. RESULTS: In this study, a multilayer attention neural graph-based collaborative filtering (MLNGCF) is proposed. MLNGCF first enhances multiple biological information with autoencoder as the initial features of circRNAs and diseases. Then, by constructing a central network of different diseases and circRNAs, a multilayer cooperative attention-based message propagation is performed on the central network to obtain the high-order features of circRNAs and diseases. A neural network-based collaborative filtering is constructed to predict the unknown circRNA-disease associations and update the model parameters. Experiments on the benchmark datasets demonstrate that MLNGCF outperforms state-of-the-art methods, and the prediction results are supported by the literature in the case studies. AVAILABILITY AND IMPLEMENTATION: The source codes and benchmark datasets of MLNGCF are available at https://github.com/ABard0/MLNGCF. Qunzhuo Wu, Zhaohong Deng, Wei Zhang 0221, Xiaoyong Pan, Kup-Sze Choi, Yun Zuo 0001, Hong-Bin Shen, Dongjun Yu |
Bioinform. | 2 |
| 2023 | Takagi-Sugeno-Kang Fuzzy System Towards Label-scarce Incomplete Multi-View Data Classification
Wei Zhang 0221, Zhaohong Deng, Qiongdan Lou, Te Zhang, Kup-Sze Choi, Shitong Wang 0001 |
Inf. Sci. | 2 |
| 2023 | EnsDeepDP: An Ensemble Deep Learning Approach for Disease Prediction Through MetagenomicsabstractA growing number of studies show that the human microbiome plays a vital role in human health and can be a crucial factor in predicting certain human diseases. However, microbiome data are often characterized by the limited samples and high-dimensional features, which pose a great challenge for machine learning methods. Therefore, this paper proposes a novel ensemble deep learning disease prediction method that combines unsupervised and supervised learning paradigms. First, unsupervised deep learning methods are used to learn the potential representation of the sample. Afterwards, the disease scoring strategy is developed based on the deep representations as the informative features for ensemble analysis. To ensure the optimal ensemble, a score selection mechanism is constructed, and performance boosting features are engaged with the original sample. Finally, the composite features are trained with gradient boosting classifier for health status decision. For case study, the ensemble deep learning flowchart has been demonstrated on six public datasets extracted from the human microbiome profiling. The results show that compared with the existing algorithms, our framework achieves better performance on disease prediction. Jinlin Zhu, Zhaohong Deng, Wenwei Lu |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2023 | End-to-End Incomplete Multiview Fuzzy Clustering With Adaptive Missing View Imputation and Cooperative LearningabstractThe purpose of multiview fuzzy clustering is to integrate fuzzy partitions of multiple complete views and obtain the optimal clustering partition. However, multiview data collected in the real world are often incomplete. To address this problem, partial view-based methods and missing view imputation-based methods have been proposed in recent years and achieved some success. However, these two groups of methods still have the following issues. First, since missing view processing and clustering are handled separately, the processed data lack relevance for clustering. Second, either hidden or visible information is explored for clustering, which precludes cooperative learning between these two types of information. Third, the within view and between view information is not fully explored. This article proposes a new end-to-end incomplete multiview fuzzy clustering method to deal with the issues. Based on traditional multiview fuzzy clustering, we construct a new end-to-end clustering framework to integrate the three tasks—missing view imputation, hidden view learning and clustering—as a single process. The framework not only enables mutual coordination of the three tasks, but also cooperative learning between the hidden and visible views. Next, we explore the within and between view information and propose two enhanced learning mechanisms to improve the quality of the imputed missing views and the learned hidden views. Finally, we introduce an adaptive view weighting mechanism to further improve the robustness of the model. Experiments on real-world datasets demonstrate the superiority of the proposed method. Wei Zhang 0221, Zhaohong Deng, Kup-Sze Choi, Shitong Wang 0001 |
IEEE Trans. Fuzzy Syst. | 2 |
| 2022 | An Efficient Scheduling Strategy for Containers Based on Kubernetes
Xurong Zhang, Yuan Liu 0021, Zhaohong Deng |
CollaborateCom (1) | 4 |
| 2022 | circRNA-binding protein site prediction based on multi-view deep learning, subspace learning and multi-view classifierabstractCircular RNAs (circRNAs) generally bind to RNA-binding proteins (RBPs) to play an important role in the regulation of autoimmune diseases. Thus, it is crucial to study the binding sites of RBPs on circRNAs. Although many methods, including traditional machine learning and deep learning, have been developed to predict the interactions between RNAs and RBPs, and most of them are focused on linear RNAs. At present, few studies have been done on the binding relationships between circRNAs and RBPs. Thus, in-depth research is urgently needed. In the existing circRNA-RBP binding site prediction methods, circRNA sequences are the main research subjects, but the relevant characteristics of circRNAs have not been fully exploited, such as the structure and composition information of circRNA sequences. Some methods have extracted different views to construct recognition models, but how to efficiently use the multi-view data to construct recognition models is still not well studied. Considering the above problems, this paper proposes a multi-view classification method called DMSK based on multi-view deep learning, subspace learning and multi-view classifier for the identification of circRNA-RBP interaction sites. In the DMSK method, first, we converted circRNA sequences into pseudo-amino acid sequences and pseudo-dipeptide components for extracting high-dimensional sequence features and component features of circRNAs, respectively. Then, the structure prediction method RNAfold was used to predict the secondary structure of the RNA sequences, and the sequence embedding model was used to extract the context-dependent features. Next, we fed the above four views' raw features to a hybrid network, which is composed of a convolutional neural network and a long short-term memory network, to obtain the deep features of circRNAs. Furthermore, we used view-weighted generalized canonical correlation analysis to extract four views' common features by subspace learning. Finally, the learned subspace common features and multi-view deep features were fed to train the downstream multi-view TSK fuzzy system to construct a fuzzy rule and fuzzy inference-based multi-view classifier. The trained classifier was used to predict the specific positions of the RBP binding sites on the circRNAs. The experiments show that the prediction performance of the proposed method DMSK has been improved compared with the existing methods. The code and dataset of this study are available at https://github.com/Rebecca3150/DMSK. Zhaohong Deng, Xiaoyong Pan, Zhisheng Wei, Hong-Bin Shen, Kup-Sze Choi, Shitong Wang 0001, Jing Wu 0030 |
Briefings Bioinform. | 2 |
| 2022 | MDGF-MCEC: a multi-view dual attention embedding model with cooperative ensemble learning for CircRNA-disease association predictionabstractCircular RNA (circRNA) is closely involved in physiological and pathological processes of many diseases. Discovering the associations between circRNAs and diseases is of great significance. Due to the high-cost to verify the circRNA-disease associations by wet-lab experiments, computational approaches for predicting the associations become a promising research direction. In this paper, we propose a method, MDGF-MCEC, based on multi-view dual attention graph convolution network (GCN) with cooperative ensemble learning to predict circRNA-disease associations. First, MDGF-MCEC constructs two disease relation graphs and two circRNA relation graphs based on different similarities. Then, the relation graphs are fed into a multi-view GCN for representation learning. In order to learn high discriminative features, a dual-attention mechanism is introduced to adjust the contribution weights, at both channel level and spatial level, of different features. Based on the learned embedding features of diseases and circRNAs, nine different feature combinations between diseases and circRNAs are treated as new multi-view data. Finally, we construct a multi-view cooperative ensemble classifier to predict the associations between circRNAs and diseases. Experiments conducted on the CircR2Disease database demonstrate that the proposed MDGF-MCEC model achieves a high area under curve of 0.9744 and outperforms the state-of-the-art methods. Promising results are also obtained from experiments on the circ2Disease and circRNADisease databases. Furthermore, the predicted associated circRNAs for hepatocellular carcinoma and gastric cancer are supported by the literature. The code and dataset of this study are available at https://github.com/ABard0/MDGF-MCEC. Qunzhuo Wu, Zhaohong Deng, Xiaoyong Pan, Hong-Bin Shen, Kup-Sze Choi, Shitong Wang 0001, Jing Wu 0030, Dongjun Yu |
Briefings Bioinform. | 2 |
| 2022 | Monotonic relation-constrained Takagi-Sugeno-Kang fuzzy system
Zhaohong Deng, Ya Cao, Qiongdan Lou, Kup-Sze Choi, Shitong Wang 0001 |
Inf. Sci. | 1 |
| 2022 | Double-coupling learning for multi-task data stream classification
Yingzhong Shi, Andong Li, Zhaohong Deng, Qisheng Yan, Qiongdan Lou, Haoran Chen 0003, Kup-Sze Choi, Shitong Wang 0001 |
Inf. Sci. | 3 |
| 2022 | Transductive Multiview Modeling With Interpretable Rules, Matrix Factorization, and Cooperative LearningabstractMultiview fuzzy systems aim to deal with fuzzy modeling in multiview scenarios effectively and to obtain the interpretable model through multiview learning. However, current studies of multiview fuzzy systems still face several challenges, one of which is how to achieve efficient collaboration between multiple views when there are few labeled data. To address this challenge, this article explores a novel transductive multiview fuzzy modeling method. The dependency on labeled data is reduced by integrating transductive learning into the fuzzy model to simultaneously learn both the model and the labels using a novel learning criterion. Matrix factorization is incorporated to further improve the performance of the fuzzy model. In addition, collaborative learning between multiple views is used to enhance the robustness of the model. The experimental results indicate that the proposed method is highly competitive with other multiview learning methods. Wei Zhang 0221, Zhaohong Deng, Jun Wang 0024, Kup-Sze Choi, Te Zhang, Xiaoqing Luo, Hong-Bin Shen, Wenhao Ying, Shitong Wang 0001 |
IEEE Trans. Cybern. | 2 |
| 2022 | Enhanced Multiview Fuzzy Clustering Using Double Visible-Hidden View Cooperation and Network LASSO ConstraintabstractMultiview clustering is an important topic in multiview learning, where the cooperation of different views is used to improve clustering performance. Although multiview clustering has made considerable progress, most existing methods only utilize the information of the original visible views, or only consider some hidden space information shared by different views. Two of the challenges are: 1) insufficient exploitation of cooperative learning between visible and hidden information despite some preliminary attempts, and 2) inadequate consideration of topological information for improving multiview clustering. To meet the challenges, we propose the cooperation enhanced multiview fuzzy clustering method (CE-MVFC) in this article. First, we characterize multiview data with two hidden views, which are obtained by adaptive multiview non-negative matrix factorization (NMF) and fuzzy partition information of each sample in different clusters. Then, we integrated the hidden views and the original visible views to realize visible-hidden cooperation learning. Furthermore, we establish a similarity matrix for each visible view and the hidden view obtained through NMF to describe the data topology in these views. Based on the spatial topological relationship of the samples and the representation of hidden view obtained by fuzzy partition, the network least absolute shrinkage and selection operator is constructed to constrain multiview learning. Finally, we develop the multiview clustering method by exploiting the visible-hidden information cooperation and the spatial topological information constraints. Experiments on benchmark multiview datasets are conducted to demonstrate the highly competitive performance of the proposed CE-MVFC against the state-of-the-art methods. Zhaohong Deng, Hongtan Yang, Wei Zhang 0221, Qiongdan Lou, Kup-Sze Choi, Te Zhang, Jin Zhou 0003, Shitong Wang 0001 |
IEEE Trans. Fuzzy Syst. | 1 |
| 2022 | Multilabel Takagi-Sugeno-Kang Fuzzy SystemabstractMultilabel (ML) classification can effectively identify the relevant labels of an instance from a given set of labels. However, the modeling of the relationship between the features and the labels is critical to classification performance. To this end, in this article, we propose a new ML classification method, called ML Takagi-Sugeno-Kang fuzzy system (ML-TSK FS), to improve the classification performance. The structure of ML-TSK FS is designed using fuzzy rules to model the relationship between features and labels. The FS is trained by integrating fuzzy inference-based ML correlation learning with ML regression loss. The proposed ML-TSK FS is evaluated experimentally on 12 benchmark ML datasets. The results show that the performance of ML-TSK FS is competitive with existing methods in terms of various evaluation metrics, indicating that it is able to model the feature-label relationship effectively using fuzzy inference rules and enhances the classification performance. Qiongdan Lou, Zhaohong Deng, Zhiyong Xiao 0001, Kup-Sze Choi, Shitong Wang 0001 |
IEEE Trans. Fuzzy Syst. | 2 |
| 2022 | Manifold-Regularized Multitask Fuzzy System Modeling With Low-Rank and Sparse Structures in Consequent ParametersabstractMultitask modeling methods for Takagi–Sugeno–Kang (TSK) fuzzy systems exhibit better generalization ability attributed to the utilization of the knowledge of intertask correlation. However, existing methods usually ignore the balance between the sharing of the common knowledge across multiple tasks and the preservation of the task-specific characteristics of each rule. To this end, we propose a novel manifold-regularized multitask modeling method for TSK fuzzy system by introducing low-rank and sparse structures into consequent parameters across multiple tasks. Specifically, we decompose the consequent parameters into two components—a task–shared component that represents similar structure across multiple tasks, and a task-specific component that encodes the sparse characteristics of the individual tasks. This can be implemented by imposing low-rank constraints on the task-shared component and applying the sparse constraints on the task-specific component. A new manifold regularization is further devised to reflect the feature-feature relation, which provides prior knowledge in multitask learning. An efficient augmented Lagrange multiplier is developed to solve the optimization problem. The experimental results demonstrate that the proposed model significantly outperforms the existing methods. Jun Wang 0024, Zhuangzhuang Zhao, Zhaohong Deng, Kup-Sze Choi, Lejun Gong, Jun Shi 0004, Shitong Wang 0001 |
IEEE Trans. Fuzzy Syst. | 3 |
| 2022 | Incomplete Multiple View Fuzzy Inference System With Missing View Imputation and Cooperative LearningabstractAdvancement of technology has made available data of different modalities that can be integrated effectively through multiple view learning for modeling real-world problems. Although multiple view learning has achieved great success in many applications, it still faces several challenges. One of them is how to reduce the negative impact of the missing views in incomplete multiple view datasets by fully exploiting the information available. Another challenge is how to enhance the interpretability of the multiple view model for scenarios with high transparency requirement. To address these challenges, this article proposes a novel modeling method for incomplete multiple view fuzzy system. Based on fuzzy interpretable rules, the method integrates missing view imputation and hidden view learning as one single process to yield a model of high interpretability, where cooperative learning is used to mine the complementary information between the visible views and the hidden view. The proposed method has four advantages when compared with existing approaches: 1) the method is more interpretable, attributed to the fuzzy interpretable rules that it is based on, 2) missing view imputation is integrated into the modeling to make it more efficient than the existing two-step strategy, 3) the method not only imputes missing views, but also mines the hidden view shared by the multiple visible views, and 4) cooperative learning is used to mine the complementary information, which significantly reduces the negative impact of missing views. Experiments on real datasets demonstrate the advantages of the proposed method. Wei Zhang 0221, Zhaohong Deng, Te Zhang, Kup-Sze Choi, Jun Wang 0024, Shitong Wang 0001 |
IEEE Trans. Fuzzy Syst. | 2 |
| 2022 | Multi-View Clustering With the Cooperation of Visible and Hidden ViewsabstractMulti-view data are becoming common in real-world applications and many multi-view clustering algorithms have thus been proposed. The existing algorithms usually focus on the cooperation of different visible views in the original space but neglect the influence of the hidden information among these visible views, or they only consider the hidden information among the views. The algorithms are therefore not efficient since the available information is not fully exploited, particularly the otherness information in different views and the consistency information among them. In practice, the otherness and consistency information in multi-view data are both very useful for effective clustering analyses. In this study, a Multi-View clustering algorithm with the Cooperation of Visible and Hidden views, i.e., MV-Co-VH, is proposed. The MV-Co-VH algorithm first projects the multiple views from different visible spaces to the common hidden space by using non-negative matrix factorization to obtain the common hidden view data. Collaborative learning is then implemented in the clustering procedure based on the visible views and the shared hidden view. The experimental results of extensive experiments on UCI multi-view datasets and real-world image multi-view datasets show that the clustering performance of the proposed algorithm is competitive with or even better than that of the existing algorithms. Zhaohong Deng, Ruixiu Liu, Peng Xu 0051, Kup-Sze Choi, Wei Zhang 0221, Xiaobin Tian, Te Zhang, Bin Qin 0003, Shitong Wang 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2021 | RNA-binding protein recognition based on multi-view deep feature and multi-label learningabstractRNA-binding protein (RBP) is a class of proteins that bind to and accompany RNAs in regulating biological processes. An RBP may have multiple target RNAs, and its aberrant expression can cause multiple diseases. Methods have been designed to predict whether a specific RBP can bind to an RNA and the position of the binding site using binary classification model. However, most of the existing methods do not take into account the binding similarity and correlation between different RBPs. While methods employing multiple labels and Long Short Term Memory Network (LSTM) are proposed to consider binding similarity between different RBPs, the accuracy remains low due to insufficient feature learning and multi-label learning on RNA sequences. In response to this challenge, the concept of RNA-RBP Binding Network (RRBN) is proposed in this paper to provide theoretical support for multi-label learning to identify RBPs that can bind to RNAs. It is experimentally shown that the RRBN information can significantly improve the prediction of unknown RNA-RBP interactions. To further improve the prediction accuracy, we present the novel computational method iDeepMV which integrates multi-view deep learning technology under the multi-label learning framework. iDeepMV first extracts data from the views of amino acid sequence and dipeptide component based on the RNA sequences as the original view. Deep neural network models are then designed for the respective views to perform deep feature learning. The extracted deep features are fed into multi-label classifiers which are trained with the RNA-RBP interaction information for the three views. Finally, a voting mechanism is designed to make comprehensive decision on the results of the multi-label classifiers. Our experimental results show that the prediction performance of iDeepMV, which combines multi-view deep feature learning models with RNA-RBP interaction information, is significantly better than that of the state-of-the-art methods. iDeepMV is freely available at http://www.csbio.sjtu.edu.cn/bioinf/iDeepMV for academic use. The code is freely available at http://github.com/uchihayht/iDeepMV. Zhaohong Deng, Xiaoyong Pan, Hong-Bin Shen, Kup-Sze Choi, Shitong Wang 0001, Jing Wu 0030 |
Briefings Bioinform. | 2 |
| 2021 | Transfer Representation Learning With TSK Fuzzy SystemabstractTransfer learning can address the learning tasks of unlabeled data in the target domain by leveraging plenty of labeled data from a different but related source domain. A core issue in transfer learning is to learn a shared feature space where the distributions of the data from the two domains are matched. This learning process can be named as transfer representation learning (TRL). Feature transformation methods are crucial to ensure the success of TRL. The most commonly used feature transformation method in TRL is kernel-based nonlinear mapping to the high-dimensional space, followed by linear dimensionality reduction. But the kernel functions are lack of interpretability, and it is difficult to select kernel functions. To this end, this article proposes a more intuitive and interpretable method, called TRL with TSK-FS (TRL-TSK-FS), by combining TSK fuzzy system (TSK-FS) with transfer learning. Specifically, TRL-TSK-FS realizes TRL from two aspects. On one hand, the data in the source and target domains are transformed into the fuzzy feature space where the distribution distance of the data between the two domains is minimized. On the other hand, discriminant information and geometric properties of the data are preserved by linear discriminant analysis and principal component analysis. A further advantage is that nonlinear transformation is realized in the proposed method by constructing fuzzy mapping with the antecedent part of the TSK-FS instead of kernel functions, which are difficult to be selected. Extensive experiments are conducted on text and image datasets to demonstrate the superiority of the proposed method. Peng Xu 0051, Zhaohong Deng, Jun Wang 0024, Qun Zhang 0004, Kup-Sze Choi, Shitong Wang 0001 |
IEEE Trans. Fuzzy Syst. | 2 |
| 2021 | Robust TSK Fuzzy System Based on Semisupervised Learning for Label Noise DataabstractAs an important branch in the field of soft computing, TSK fuzzy systems have been diversely applied to supervised learning in recent years. However, real-world data may contain label noise, which has a negative impact on supervised learning. Label noise samples change the distribution of samples in each class, mislead learning algorithms and make classification problems more complicated. There are various sources of label noise, such as wrong assignment of labels during the data collection, contamination during the data storage, and so on. Thus, it is usually costly and time-consuming to obtain data with no label noise. When dealing with label noise data, existing TSK fuzzy system algorithms still have room for improvement. This article proposes a robust TSK fuzzy system based on semisupervised learning for label noise data (RTSK-FS-SS). By introducing an intuitionistic fuzzy set method, the proposed algorithm can detect label noise samples. An improved learning vector quantization is further adopted to overcome the challenge that traditional unsupervised learning-based antecedent part generation processes unable to make full use of the label information of training samples. Finally, we discard the label of suspicious samples and a consequent parameter learning method based on semisupervised learning is proposed. The proposed algorithm is validated using extensive experiments. Te Zhang, Zhaohong Deng, Hisao Ishibuchi, Lie Meng Pang |
IEEE Trans. Fuzzy Syst. | 2 |
| 2021 | Multitask TSK Fuzzy System Modeling by Jointly Reducing Rules and Consequent ParametersabstractExisting multitask Takagi-Sugeno-Kang (TSK) fuzzy modeling methods always produce high complex fuzzy models with numerous redundant rules and consequent parameters. To this end, we propose a novel multitask TSK fuzzy modeling method called mtSparseTSK, which learns a compact set of fuzzy rules and shared consequent parameters across tasks in a unified procedure. Specifically, we consider the fuzzy rule reduction and consequent parameter selection across tasks by devising novel group sparsity regularizations in the learning criterion of the model. We also integrate the intertask relations in the proposed TSK model for multitask learning. We fully utilize the block structure in the TSK fuzzy models in formulating a joint block sparse optimization problem and develop a procedure for alternating direction method of multipliers (ADMMs) to find the optimal solution of the problem. Experiments on the synthetic and real-world datasets demonstrate the distinctive performance of the proposed methods over the existing ones on multitask fuzzy system modeling. Jun Wang 0024, Zhaohong Deng, Yizhang Jiang, Jihua Zhu, Lei Chen 0011, Lejun Gong, Shitong Wang 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2019 | Multi-View Information-Theoretic Co-Clustering for Co-Occurrence DataabstractMulti-view clustering has received much attention recently. Most of the existing multi-view clustering methods only focus on one-sided clustering. As the co-occurring data elements involve the counts of sample-feature co-occurrences, it is more efficient to conduct two-sided clustering along the samples and features simultaneously. To take advantage of two-sided clustering for the co-occurrences in the scene of multi-view clustering, a two-sided multi-view clustering method is proposed, i.e., multi-view information-theoretic co-clustering (MV-ITCC). The proposed method realizes two-sided clustering for co-occurring multi-view data under the formulation of information theory. More specifically, it exploits the agreement and disagreement among views by sharing a common clustering results along the sample dimension and keeping the clustering results of each view specific along the feature dimension. In addition, the mechanism of maximum entropy is also adopted to control the importance of different views, which can give a right balance in leveraging the agreement and disagreement. Extensive experiments are conducted on text and image multiview datasets. The results clearly demonstrate the superiority of the proposed method. Peng Xu 0051, Zhaohong Deng, Kup-Sze Choi, Longbing Cao, Shitong Wang 0001 |
AAAI | 2 |
| 2019 | Interpretable Feature Learning Using Multi-output Takagi-Sugeno-Kang Fuzzy System for Multi-center ASD Diagnosis
Jun Wang 0024, Tao Zhou 0002, Zhaohong Deng, Huifang Huang, Shitong Wang 0001, Jun Shi 0004, Dinggang Shen |
MICCAI (3) | 4 |
| 2019 | Generalized Hidden-Mapping Transductive Transfer Learning for Recognition of Epileptic Electroencephalogram SignalsabstractElectroencephalogram (EEG) signal identification based on intelligent models is an important means in epilepsy detection. In the recognition of epileptic EEG signals, traditional intelligent methods usually assume that the training dataset and testing dataset have the same distribution, and the data available for training are adequate. However, these two conditions cannot always be met in practice, which reduces the ability of the intelligent recognition model obtained in detecting epileptic EEG signals. To overcome this issue, an effective strategy is to introduce transfer learning in the construction of the intelligent models, where knowledge is learned from the related scenes (source domains) to enhance the performance of model trained in the current scene (target domain). Although transfer learning has been used in EEG signal identification, many existing transfer learning techniques are designed only for a specific intelligent model, which limit their applicability to other classical intelligent models. To extend the scope of application, the generalized hidden-mapping transductive learning method is proposed to realize transfer learning for several classical intelligent models, including feedforward neural networks, fuzzy systems, and kernelized linear models. These intelligent models can be trained effectively by the proposed method even though the data available are insufficient for model training, and the generalization abilities of the trained model is also enhanced by transductive learning. A number of experiments are carried out to demonstrate the effectiveness of the proposed method in epileptic EEG recognition. The results show that the method is highly competitive or superior to some existing state-of-the-art methods. Lixiao Xie, Zhaohong Deng, Peng Xu 0051, Kup-Sze Choi, Shitong Wang 0001 |
IEEE Trans. Cybern. | 2 |
| 2019 | Concise Fuzzy System Modeling Integrating Soft Subspace Clustering and Sparse LearningabstractThe superior interpretability and uncertainty modeling ability of Takagi-Sugeno-Kang fuzzy system (TSK FS) make it possible to describe complex nonlinear systems intuitively and efficiently. However, classical TSK FS usually adopts the whole feature space of the data for model construction, which can result in lengthy rules for high-dimensional data and lead to degeneration in interpretability. Furthermore, for highly nonlinear modeling task, it is usually necessary to use a large number of rules which further weaken the clarity and interpretability of TSK FS. To address these issues, an enhanced soft subspace clustering (ESSC) and sparse learning (SL) based concise zero-order TSK FS construction method, called ESSC-SL-CTSK-FS, is proposed in this paper by integrating the techniques of ESSC and SL. In this method, ESSC is used to generate the antecedents and various sparse subspaces for different fuzzy rules, whereas SL is used to optimize the consequent parameters of the fuzzy rules based on which the number of fuzzy rules can be effectively reduced. Finally, the proposed ESSC-SL-CTSK-FS method is used to construct concise zero-order TSK FS that can explain the scenes in high-dimensional data modeling more clearly and easily. Experiments are conducted on various real-world datasets to confirm the advantages. Peng Xu 0051, Zhaohong Deng, Te Zhang, Kup-Sze Choi, Suhang Gu, Jun Wang 0024, Shitong Wang 0001 |
IEEE Trans. Fuzzy Syst. | 2 |
| 2019 | Multiview Fuzzy Logic System With the Cooperation Between Visible and Hidden ViewsabstractMultiview datasets are frequently encountered in learning tasks, such as web data mining and multimedia information analysis. Given a multiview dataset, traditional learning algorithms usually decompose it into several single-view datasets, from each of which a single-view model is learned. In contrast, a multiview learning algorithm can achieve better performance by cooperative learning on the multiview data. However, existing multiview approaches mainly focus on the views that are visible and ignore the hidden information behind the visible views, which usually contains some intrinsic information of the multiview data, or vice versa. To address this problem, this paper proposes a multiview fuzzy logic system which utilizes both the hidden information shared by the multiple visible views and the information of each visible view. Extensive experiments were conducted to validate its effectiveness. Te Zhang, Zhaohong Deng, Dongrui Wu, Shitong Wang 0001 |
IEEE Trans. Fuzzy Syst. | 2 |
| 2018 | A Novel Takagi-Sugeno Fuzzy System Modeling Method with Joint Feature Selection and Rule ReductionabstractTraditional Takagi-Sugeno (T-S) fuzzy system modeling methods always yield a large number of fuzzy rules. Besides, they also include almost all the original features in the final model. These two factors make the final model sophisticated. In this paper, we propose a novel T-S fuzzy system modeling method called GS-FIS (Group Sparse Fuzzy Inference Systems), which performs fuzzy rule reduction and feature selection simultaneously in a unified framework. Considering the group structure information in the T-S fuzzy system and common features among fuzzy rules, we cast the fuzzy system modeling into a joint group sparse optimization problem and further develop an alternating direction method of multipliers procedure to derive the optimum solution to the problem. Experimental results on the synthetic dataset and several real-world datasets show that the proposed method can not only obtain a satisfactory generalization performance but also reduce the number of fuzzy rules and features effectively. Jun Wang 0024, Jihua Zhu, Yizhang Jiang, Zhaohong Deng, Weiwei Li 0001, Shitong Wang 0001 |
FUZZ-IEEE | 6 |
| 2018 | Generalized Hidden-Mapping Minimax Probability Machine for the training and reliability learning of several classical intelligent models
Zhaohong Deng, Junyong Chen, Te Zhang, Longbing Cao, Shitong Wang 0001 |
Inf. Sci. | 1 |
| 2018 | Cascaded Hidden Space Feature Mapping, Fuzzy Clustering, and Nonlinear Switching Regression on Large DatasetsabstractThe success of fuzzy clustering heavily relies on the features of the input data. Based on the fact that deep architectures are able to more accurately characterize the data representations in a layer-by-layer manner, this paper proposes a novel feature mapping technique called cascaded hidden-space (CHS) feature mapping and investigates its combination with classical fuzzy c-means (FCM) and fuzzy c-regressions (FCR). Since the parameters between the layers of CHS feature mapping are randomly generated and need not be tuned layer-by-layer, CHS is easily implemented with less training data. By performing classical FCM in CHS, a novel fuzzy clustering framework called CHS-FCM is developed; several of its variants are presented using different dimension-reduction methods in a CHS-FCM clustering framework. The combination of CHS-FCM with nonlinear switch regressions is called CHS-FCR, and it performs FCR in CHS. The proposed CHS-FCR provides better results than FCR for nonlinear process modeling. Both CHS-FCM and CHS-FCR exhibit low memory consumption and require less training data. The experimental results verify the superiority of the proposed methods over classical fuzzy clustering methods. Jun Wang 0024, Xiaohua Qian, Yizhang Jiang, Zhaohong Deng, Shitong Wang 0001 |
IEEE Trans. Fuzzy Syst. | 5 |
| 2018 | Data-Driven Elastic Fuzzy Logic System Modeling: Constructing a Concise System With Human-Like Inference MechanismabstractThe construction of fuzzy logic systems (FLSs) using data-driven techniques has become the most popular modeling approach. However, this approach still faces critical challenges, including the difficulty in obtaining concise models for high-dimensional data and generating accurate fuzzy rules to simulate human inference mechanism. To tackle these issues, a new FLS modeling framework called data-driven elastic FLS (DD-EFLS) is proposed in this paper. The DD-EFLS has two key characteristics. First, the fuzzy rules in the rule base can use different feature subspaces that are extracted from the original high-dimensional space to yield simple and accurate rules in feature spaces of lower dimensionality. Second, fuzzy inferences from various views are implemented by embedding different rules in the corresponding subspaces to imitate human inference mechanism. Based on the DD-EFLS framework, an elastic Takagi-Sugeno-Kang (TSK) FLS modeling method (ETSK-FLS) is proposed to train the elastic TSK FLS using the concise rules and a more human-like inference mechanism for modeling tasks based on high-dimensional datasets. The characteristics and advantages of the proposed framework and the ETSK-FLS method are validated experimentally using both synthetic and real-world datasets. Jiangbin Zhang, Zhaohong Deng, Kup-Sze Choi, Shitong Wang 0001 |
IEEE Trans. Fuzzy Syst. | 2 |
| 2018 | Tackling Missing Data in Community Health Studies Using Additive LS-SVM ClassifierabstractMissing data is a common issue in community health and epidemiological studies. Direct removal of samples with missing data can lead to reduced sample size and information bias, which deteriorates the significance of the results. While data imputation methods are available to deal with missing data, they are limited in performance and could introduce noises into the dataset. Instead of data imputation, a novel method based on additive least square support vector machine (LS-SVM) is proposed in this paper for predictive modeling when the input features of the model contain missing data. The method also determines simultaneously the influence of the features with missing values on the classification accuracy using the fast leave-one-out cross-validation strategy. The performance of the method is evaluated by applying it to predict the quality of life (QOL) of elderly people using health data collected in the community. The dataset involves demographics, socioeconomic status, health history, and the outcomes of health assessments of 444 community-dwelling elderly people, with 5% to 60% of data missing in some of the input features. The QOL is measured using a standard questionnaire of the World Health Organization. Results show that the proposed method outperforms four conventional methods for handling missing data-case deletion, feature deletion, mean imputation, and K-nearest neighbor imputation, with the average QOL prediction accuracy reaching 0.7418. It is potentially a promising technique for tackling missing data in community health research and other applications. Guanjin Wang, Zhaohong Deng, Kup-Sze Choi |
IEEE J. Biomed. Health Informatics | 2 |
| 2017 | Robust extreme learning fuzzy systems using ridge regression for small and noisy datasetsabstractFuzzy Extreme Learning Machine (F-ELM) constructs a fuzzy neural networks by embedding fuzzy membership functions and rules into the hidden layer of extreme learning machine (ELM), that is, it can be interpreted as a fuzzy system with the structure of neural network. Although F-ELM has shown the characteristics of fast learning of model parameters, it has poor robustness to small and noisy datasets since its parameters connecting hidden layer with output layer are optimized by least square(LS). In order to overcome this challenge, a Ridge Regression based Extreme Learning Fuzzy System (RR-EL-FS) is presented in this study, which has introduced the strategy of ridge regression into F-ELM to enhance the robustness. The experimental results also validate that the performance of RR-EL-FS is better than F-ELM and some related methods to small and noisy datasets. Te Zhang, Zhaohong Deng, Kup-Sze Choi, Jiefang Liu, Shitong Wang 0001 |
FUZZ-IEEE | 2 |
| 2017 | Detection of epilepsy with Electroencephalogram using rule-based classifiers
Guanjin Wang, Zhaohong Deng, Kup-Sze Choi |
Neurocomputing | 2 |
| 2017 | Knowledge-leveraged transfer fuzzy C-Means for texture image segmentation with self-adaptive cluster prototype matching
Pengjiang Qian, Kaifa Zhao, Yizhang Jiang, Kuan-Hao Su, Zhaohong Deng, Shitong Wang 0001, Raymond F. Muzic Jr. |
Knowl. Based Syst. | 5 |
| 2017 | Recognition of Epileptic EEG Signals Using a Novel Multiview TSK Fuzzy SystemabstractRecognition of epileptic electroencephalogram (EEG) signals using machine learning techniques is becoming popular. In general, the construction of intelligent epileptic EEG recognition system involves two steps. First, an appropriate feature extraction method is applied to obtain representative features from the original raw EEG signals. Second, an effective intelligent model is trained based on the extracted features. However, there exist two major challenges in the process: 1) it is nontrivial to determine the appropriate feature extraction method to be used; 2) although many classical machine learning methods have been used for epileptic EEG recognition, most of them are “black box” approaches and more interpretable methods are desirable. To address these two challenges, a new epileptic EEG recognition method based on a multiview learning framework and fuzzy system modeling is proposed. First, multiview EEG data are generated by employing different feature extraction methods to obtain the features from different views of the signals. Second, the classical Takagi-Sugeno-Kang fuzzy system (TSK-FS) is introduced as an easy-to-interpret recognition model to develop a multiview TSK-FS method, called MV-TSK-FS, to identify epileptic EEG signals. For the proposed MV-TSK-FS, the importance of each view, i.e., the importance of each feature extraction method, can be evaluated according to the weighting of each view, and consequently the final decision can be made based on the weighted outputs of different views. Experimental results indicate that the MV-TSK-FS is a promising method when compared with the state-of-the-art algorithms. Yizhang Jiang, Zhaohong Deng, Korris Fu-Lai Chung, Guanjin Wang, Pengjiang Qian, Kup-Sze Choi, Shitong Wang 0001 |
IEEE Trans. Fuzzy Syst. | 2 |
| 2017 | Realizing Two-View TSK Fuzzy Classification System by Using Collaborative LearningabstractIn this paper, a novel Takagi-Sugeno-Kang (TSK) fuzzy classification system (FCS) is firstly presented for pattern classification tasks. It is distinguished by having the large margin criterion properly integrated into its objective function. In order to exploit the applicability of fuzzy systems in multiview scenarios, the proposed TSK-FCS is extended to a two-view version, called two-view TSK-FCS (TwoV-TSK-FCS), by using a collaborative learning mechanism. The adopted collaborative learning mechanism not only fully considers the independent information of each view, but also effectively discovers the correlation information hidden in the two views. Thus, the performance of TwoV-TSK-FCS can be enhanced accordingly. Comprehensive experiments on two-view synthetic and UCI datasets demonstrate the effectiveness of the proposed two-view FCS. Yizhang Jiang, Zhaohong Deng, Korris Fu-Lai Chung, Shitong Wang 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2016 | A survey on soft subspace clustering
Zhaohong Deng, Kup-Sze Choi, Yizhang Jiang, Jun Wang 0024, Shitong Wang 0001 |
Inf. Sci. | 1 |
| 2016 | A novel multi-task TSK fuzzy classifier and its enhanced version for labeling-risk-aware multi-task classification
Yizhang Jiang, Zhaohong Deng, Kup-Sze Choi, Korris Fu-Lai Chung, Shitong Wang 0001 |
Inf. Sci. | 2 |
| 2016 | Scalable learning method for feedforward neural networks using minimal-enclosing-ball approximation
Jun Wang 0024, Zhaohong Deng, Xiaoqing Luo, Yizhang Jiang, Shitong Wang 0001 |
Neural Networks | 2 |
| 2016 | Distance metric learning for soft subspace clustering in composite kernel space
Jun Wang 0024, Zhaohong Deng, Kup-Sze Choi, Yizhang Jiang, Xiaoqing Luo, Korris Fu-Lai Chung, Shitong Wang 0001 |
Pattern Recognit. | 2 |
| 2016 | Semi-Supervised SVM With Extended Hidden FeaturesabstractMany traditional semi-supervised learning algorithms not only train on the labeled samples but also incorporate the unlabeled samples in the training sets through an automated labeling process such as manifold preserving. If some labeled samples are falsely labeled, the automated labeling process will generally propagate negative impact on the classifier in quite a serious manner. In order to avoid such an error propagating effect, the unlabeled samples should not be directly incorporated into the training sets during the automated labeling strategy. In this paper, a new semi-supervised support vector machine with extended hidden features (SSVM-EHF) is presented to address this issue. According to the maximum margin principle and the minimum integrated squared error between the probability distributions of the labeled and unlabeled samples, the dimensionality of the labeled and unlabeled samples is extended through an orthonormal transformation to generate the corresponding hidden features shared by the labeled and unlabeled samples. After doing so, the last step in the process of training of SSVM-EHF is done only on the labeled samples with their original and hidden features, and the unlabeled samples are no longer explicitly used. Experimental results confirm the effectiveness of the proposed method. Aimei Dong, Korris Fu-Lai Chung, Zhaohong Deng, Shitong Wang 0001 |
IEEE Trans. Cybern. | 3 |
| 2016 | Cluster Prototypes and Fuzzy Memberships Jointly Leveraged Cross-Domain Maximum Entropy ClusteringabstractThe classical maximum entropy clustering (MEC) algorithm usually cannot achieve satisfactory results in the situations where the data is insufficient, incomplete, or distorted. To address this problem, inspired by transfer learning, the specific cluster prototypes and fuzzy memberships jointly leveraged (CPM-JL) framework for cross-domain MEC (CDMEC) is firstly devised in this paper, and then the corresponding algorithm referred to as CPM-JL-CDMEC and the dedicated validity index named fuzzy memberships-based cross-domain difference measurement (FM-CDDM) are concurrently proposed. In general, the contributions of this paper are fourfold: 1) benefiting from the delicate CPM-JL framework, CPM-JL-CDMEC features high-clustering effectiveness and robustness even in some complex data situations; 2) the reliability of FM-CDDM has been demonstrated to be close to well-established external criteria, e.g., normalized mutual information and rand index, and it does not require additional label information. Hence, using FM-CDDM as a dedicated validity index significantly enhances the applicability of CPM-JL-CDMEC under realistic scenarios; 3) the performance of CPM-JL-CDMEC is generally better than, at least equal to, that of MEC because CPM-JL-CDMEC can degenerate into the standard MEC algorithm after adopting the proper parameters, and which avoids the issue of negative transfer; and 4) in order to maximize privacy protection, CPM-JL-CDMEC employs the known cluster prototypes and their associated fuzzy memberships rather than the raw data in the source domain as prior knowledge. The experimental studies thoroughly evaluated and demonstrated these advantages on both synthetic and real-life transfer datasets. Pengjiang Qian, Yizhang Jiang, Zhaohong Deng, Lingzhi Hu, Shouwei Sun, Shitong Wang 0001, Raymond F. Muzic Jr. |
IEEE Trans. Cybern. | 3 |
| 2016 | Transfer Prototype-Based Fuzzy ClusteringabstractTraditional prototype-based clustering methods, such as the well-known fuzzy c-means (FCM) algorithm, usually need sufficient data to find a good clustering partition. If available data are limited or scarce, most of them are no longer effective. While the data for the current clustering task may be scarce, there is usually some useful knowledge available in the related scenes/domains. In this study, the concept of transfer learning is applied to prototype-based fuzzy clustering (PFC). Specifically, the idea of leveraging knowledge from the source domain is exploited to develop a set of transfer PFC algorithms. First, two representative PFC algorithms, namely, FCM and fuzzy subspace clustering, have been chosen to incorporate with knowledge leveraging mechanisms to develop the corresponding transfer clustering algorithms based on an assumption that there are the same number of clusters between the target domain (current scene) and the source domain (related scene). Furthermore, two extended versions are also proposed to implement the transfer learning for the situation that there are different numbers of clusters between two domains. The novel objective functions are proposed to integrate the knowledge from the source domain with the data in the target domain for the clustering in the target domain. The proposed algorithms have been validated on different synthetic and real-world datasets. Experimental results demonstrate their effectiveness in comparison with both the original PFC algorithms and the related clustering algorithms like multitask clustering and coclustering. Zhaohong Deng, Yizhang Jiang, Korris Fu-Lai Chung, Hisao Ishibuchi, Kup-Sze Choi, Shitong Wang 0001 |
IEEE Trans. Fuzzy Syst. | 1 |
| 2016 | Takagi-Sugeno-Kang Transfer Learning Fuzzy Logic System for the Adaptive Recognition of Epileptic Electroencephalogram SignalsabstractThe intelligent recognition of electroencephalogram (EEG) signals has become an important approach to the detection of epilepsy. Among existing intelligent identification methods, fuzzy logic systems (FLSs) have shown a distinctive advantage in identifying epileptic EEG signals because of their strong learning abilities and interpretability. Like many conventional intelligent methods for recognizing EEG signals, in the training of FLS, it is assumed that the training dataset and test dataset are drawn from data that are identically distributed. However, this assumption is not necessarily valid in practice as it is not uncommon for the two datasets to have different distributions. To overcome this problem, a strategy is presented in this paper to construct a Takagi-Sugeno-Kang (TSK) FLS based on transductive transfer learning for identifying epileptic EEG signals. Two novel objective functions, achieved by integrating the transductive transfer learning mechanism, are proposed for the training of the TSK FLS. As regression and binary classification are two common approaches to multiclass classification, the TSK transfer learning FLS algorithms for regression and binary classification are developed, respectively, to construct the corresponding TSK FLS. Both algorithms are further used to perform a multiclass classification to recognize epileptic EEG signals. Their performance in the epileptic EEG datasets indicates promise in dealing with situations where the training and test datasets differ with regard to data distribution. Changjian Yang, Zhaohong Deng, Kup-Sze Choi, Shitong Wang 0001 |
IEEE Trans. Fuzzy Syst. | 2 |
| 2016 | Enhanced Knowledge-Leverage-Based TSK Fuzzy System Modeling for Inductive Transfer LearningabstractThe knowledge-leverage-based Takagi--Sugeno--Kang fuzzy system (KL-TSK-FS) modeling method has shown promising performance for fuzzy modeling tasks where transfer learning is required. However, the knowledge-leverage mechanism of the KL-TSK-FS can be further improved. This is because available training data in the target domain are not utilized for the learning of antecedents and the knowledge transfer mechanism from a source domain to the target domain is still too simple for the learning of consequents when a Takagi--Sugeno--Kang fuzzy system (TSK-FS) model is trained in the target domain. The proposed method, that is, the enhanced KL-TSK-FS (EKL-TSK-FS), has two knowledge-leverage strategies for enhancing the parameter learning of the TSK-FS model for the target domain using available information from the source domain. One strategy is used for the learning of antecedent parameters, while the other is for consequent parameters. It is demonstrated that the proposed EKL-TSK-FS has higher transfer learning abilities than the KL-TSK-FS. In addition, the EKL-TSK-FS has been further extended for the scene of the multisource domain. Zhaohong Deng, Yizhang Jiang, Hisao Ishibuchi, Kup-Sze Choi, Shitong Wang 0001 |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2015 | Detection of Epileptic Seizures in EEG Signals with Rule-Based Interpretation by Random Forest Approach
Guanjin Wang, Zhaohong Deng, Kup-Sze Choi |
ICIC (3) | 2 |
| 2015 | Multi-task TSK fuzzy system modeling using inter-task correlation information
Yizhang Jiang, Zhaohong Deng, Korris Fu-Lai Chung, Shitong Wang 0001 |
Inf. Sci. | 2 |
| 2015 | Multitask TSK Fuzzy System Modeling by Mining Intertask Common Hidden StructureabstractThe classical fuzzy system modeling methods implicitly assume data generated from a single task, which is essentially not in accordance with many practical scenarios where data can be acquired from the perspective of multiple tasks. Although one can build an individual fuzzy system model for each task, the result indeed tells us that the individual modeling approach will get poor generalization ability due to ignoring the intertask hidden correlation. In order to circumvent this shortcoming, we consider a general framework for preserving the independent information among different tasks and mining hidden correlation information among all tasks in multitask fuzzy modeling. In this framework, a low-dimensional subspace (structure) is assumed to be shared among all tasks and hence be the hidden correlation information among all tasks. Under this framework, a multitask Takagi-Sugeno-Kang (TSK) fuzzy system model called MTCS-TSK-FS (TSK-FS for multiple tasks with common hidden structure), based on the classical L2-norm TSK fuzzy system, is proposed in this paper. The proposed model can not only take advantage of independent sample information from the original space for each task, but also effectively use the intertask common hidden structure among multiple tasks to enhance the generalization performance of the built fuzzy systems. Experiments on synthetic and real-world datasets demonstrate the applicability and distinctive performance of the proposed multitask fuzzy system model in multitask regression learning scenarios. Yizhang Jiang, Korris Fu-Lai Chung, Hisao Ishibuchi, Zhaohong Deng, Shitong Wang 0001 |
IEEE Trans. Cybern. | 4 |
| 2015 | Collaborative Fuzzy Clustering From Multiple Weighted ViewsabstractClustering with multiview data is becoming a hot topic in data mining, pattern recognition, and machine learning. In order to realize an effective multiview clustering, two issues must be addressed, namely, how to combine the clustering result from each view and how to identify the importance of each view. In this paper, based on a newly proposed objective function which explicitly incorporates two penalty terms, a basic multiview fuzzy clustering algorithm, called collaborative fuzzy c-means (Co-FCM), is firstly proposed. It is then extended into its weighted view version, called weighted view collaborative fuzzy c-means (WV-Co-FCM), by identifying the importance of each view. The WV-Co-FCM algorithm indeed tackles the above two issues simultaneously. Its relationship with the latest multiview fuzzy clustering algorithm Collaborative Fuzzy K-Means (Co-FKM) is also revealed. Extensive experimental results on various multiview datasets indicate that the proposed WV-Co-FCM algorithm outperforms or is at least comparable to the existing state-of-the-art multitask and multiview clustering algorithms and the importance of different views of the datasets can be effectively identified. Yizhang Jiang, Korris Fu-Lai Chung, Shitong Wang 0001, Zhaohong Deng, Jun Wang 0024, Pengjiang Qian |
IEEE Trans. Cybern. | 4 |
| 2015 | Minimax Probability TSK Fuzzy System Classifier: A More Transparent and Highly Interpretable Classification ModelabstractWhen an intelligent model is used for medical diagnosis, it is desirable to have a high level of interpretability and transparent model reliability for users. Compared with most of the existing intelligence models, fuzzy systems have shown a distinctive advantage in their interpretabilities. However, how to determine the model reliability of a fuzzy system trained for a recognition task is still an unsolved problem at present. In this study, a minimax probability Takagi-Sugeno-Kang (TSK) fuzzy system classifier called MP-TSK-FSC is proposed to train a fuzzy system classifier and determine the model reliability simultaneously. For the proposed MP-TSK-FSC, a lower bound of correct classification can be presented to the users to characterize the reliability of the trained fuzzy classifier. Thus, the obtained classifier has the distinctive characteristics of both a high level of interpretability and transparent model reliability inherited from the fuzzy system and minimax probability learning strategy, respectively. Our experiments on synthetic datasets and several real-world datasets for medical diagnosis have confirmed the distinctive characteristics of the proposed method. Zhaohong Deng, Longbing Cao, Yizhang Jiang, Shitong Wang 0001 |
IEEE Trans. Fuzzy Syst. | 1 |
| 2014 | Knowledge-leverage based TSK fuzzy system with improved knowledge transferabstractIn this study, the improved knowledge-leverage based TSK fuzzy system modeling method is proposed in order to overcome the weaknesses of the knowledge-leverage based TSK fuzzy system (TSK-FS) modeling method. In particular, two improved knowledge-leverage strategies have been introduced for the parameter learning of the antecedents and consequents of the TSK-FS constructed in the current scene by transfer learning from the reference scene, respectively. With the improved knowledge-leverage learning abilities, the proposed method has shown the more adaptive modeling effect compared with traditional TSK fuzzy modeling methods and some related methods on the synthetic and real world datasets. Zhaohong Deng, Yizhang Jiang, Longbing Cao, Shitong Wang 0001 |
FUZZ-IEEE | 1 |
| 2014 | Multiple-kernel based soft subspace fuzzy clusteringabstractSoft subspace fuzzy clustering algorithms have been successfully utilized for high dimensional data in recent studies. However, the existing works often utilize only one distance function to evaluate the similarity between data items along with each feature, which leads to performance degradation for some complex data sets. In this work, a novel soft subspace fuzzy clustering algorithm MKEWFC-K is proposed by extending the existing entropy weight soft subspace clustering algorithm with a multiple-kernel learning setting. By incorporating multiple-kernel learning strategy into the framework of soft subspace fuzzy clustering, MKEWFC-K can learning the distance function adaptively during the clustering process. Moreover, it is more immune to ineffective kernels and irrelevant features in soft subspace, which makes the choice of kernels less crucial. Experiments on real-world data demonstrate the effectiveness of the proposed MKEWFC-K algorithm. Jun Wang 0024, Zhaohong Deng, Yizhang Jiang, Pengjiang Qian, Shitong Wang 0001 |
FUZZ-IEEE | 2 |
| 2014 | Transductive domain adaptive learning for epileptic electroencephalogram recognition
Changjian Yang, Zhaohong Deng, Kup-Sze Choi, Yizhang Jiang, Shitong Wang 0001 |
Artif. Intell. Medicine | 2 |
| 2014 | Double indices-induced FCM clustering and its integration with fuzzy subspace clustering
Jun Wang 0024, Korris Fu-Lai Chung, Shitong Wang 0001, Zhaohong Deng |
Pattern Anal. Appl. | 4 |
| 2014 | Generalized Hidden-Mapping Ridge Regression, Knowledge-Leveraged Inductive Transfer Learning for Neural Networks, Fuzzy Systems and Kernel MethodsabstractInductive transfer learning has attracted increasing attention for the training of effective model in the target domain by leveraging the information in the source domain. However, most transfer learning methods are developed for a specific model, such as the commonly used support vector machine, which makes the methods applicable only to the adopted models. In this regard, the generalized hidden-mapping ridge regression (GHRR) method is introduced in order to train various types of classical intelligence models, including neural networks, fuzzy logical systems and kernel methods. Furthermore, the knowledge-leverage based transfer learning mechanism is integrated with GHRR to realize the inductive transfer learning method called transfer GHRR (TGHRR). Since the information from the induced knowledge is much clearer and more concise than that from the data in the source domain, it is more convenient to control and balance the similarity and difference of data distributions between the source and target domains. The proposed GHRR and TGHRR algorithms have been evaluated experimentally by performing regression and classification on synthetic and real world datasets. The results demonstrate that the performance of TGHRR is competitive with or even superior to existing state-of-the-art inductive transfer learning algorithms. Zhaohong Deng, Kup-Sze Choi, Yizhang Jiang, Shitong Wang 0001 |
IEEE Trans. Cybern. | 1 |
| 2014 | T2FELA: Type-2 Fuzzy Extreme Learning Algorithm for Fast Training of Interval Type-2 TSK Fuzzy Logic SystemabstractA challenge in modeling type-2 fuzzy logic systems is the development of efficient learning algorithms to cope with the ever increasing size of real-world data sets. In this paper, the extreme learning strategy is introduced to develop a fast training algorithm for interval type-2 Takagi-Sugeno-Kang fuzzy logic systems. The proposed algorithm, called type-2 fuzzy extreme learning algorithm (T2FELA), has two distinctive characteristics. First, the parameters of the antecedents are randomly generated and parameters of the consequents are obtained by a fast learning method according to the extreme learning mechanism. In addition, because the obtained parameters are optimal in the sense of minimizing the norm, the resulting fuzzy systems exhibit better generalization performance. The experimental results clearly demonstrate that the training speed of the proposed T2FELA algorithm is superior to that of the existing state-of-the-art algorithms. The proposed algorithm also shows competitive performance in generalization abilities. Zhaohong Deng, Kup-Sze Choi, Longbing Cao, Shitong Wang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2013 | Weighted spherical 1-mean with phase shift and its application in electrocardiogram discord detection
Jun Wang 0024, Korris Fu-Lai Chung, Zhaohong Deng, Shitong Wang 0001, Wenhao Ying |
Artif. Intell. Medicine | 3 |
| 2013 | Fuzzy partition based soft subspace clustering and its applications in high dimensional data
Jun Wang 0024, Shitong Wang 0001, Korris Fu-Lai Chung, Zhaohong Deng |
Inf. Sci. | 4 |
| 2013 | Knowledge-Leverage-Based Fuzzy System and Its ModelingabstractThe classical fuzzy system modeling methods only consider the current scene where the training data are assumed fully collectable. However, if the available data from that scene are insufficient, the fuzzy systems trained will suffer from weak generalization for the modeling task in this scene. In order to overcome this problem, a fuzzy system with knowledge-leverage capability, which is known as a knowledge-leverage-based fuzzy system (KL-FS), is proposed in this paper. The KL-FS not only makes full use of the data from the current scene in the learning procedure but can effectively make leverage on the existing knowledge from the reference scene, e.g., the parameters of a fuzzy system obtained from a reference scene, as well. Specifically, a knowledge-leverage-based Mamdani-Larsen-type fuzzy system (KL-ML-FS) is proposed by using the reduced set density estimation technique integrating with the corresponding knowledge-leverage mechanism. The new fuzzy system modeling technique has been verified by experiments on synthetic and real-world datasets, where KL-ML-FS has better performance and adaptability than the traditional fuzzy modeling methods in scenarios with insufficient data. Zhaohong Deng, Yizhang Jiang, Korris Fu-Lai Chung, Hisao Ishibuchi, Shitong Wang 0001 |
IEEE Trans. Fuzzy Syst. | 1 |
| 2013 | Knowledge-Leverage-Based TSK Fuzzy System ModelingabstractClassical fuzzy system modeling methods consider only the current scene where the training data are assumed to be fully collectable. However, if the data available from the current scene are insufficient, the fuzzy systems trained by using the incomplete datasets will suffer from weak generalization capability for the prediction in the scene. In order to overcome this problem, a knowledge-leverage-based fuzzy system (KL-FS) is studied in this paper from the perspective of transfer learning. The KL-FS intends to not only make full use of the data from the current scene in the learning procedure, but also effectively leverage the existing knowledge from the reference scenes. Specifically, a knowledge-leverage-based Takagi-Sugeno-Kang-type Fuzzy System (KL-TSK-FS) is proposed by integrating the corresponding knowledge-leverage mechanism. The new fuzzy system modeling technique is evaluated through experiments on synthetic and real-world datasets. The results demonstrate that KL-TSK-FS has better performance and adaptability than the traditional fuzzy modeling methods in scenes with insufficient data. Zhaohong Deng, Yizhang Jiang, Kup-Sze Choi, Korris Fu-Lai Chung, Shitong Wang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2012 | Double indices induced FCM clustering and its integration with fuzzy subspace clusteringabstractFuzzy c-means is one of the most popular algorithms for clustering analysis. In this study, a novel FCM based algorithm called double indices induced FCM (DI-FCM) is developed from a new perspective. DI-FCM introduces a power exponent r into the constraints of the objective function such that the range of the fuzziness index m is extended. Furthermore, it can be explained from the perspective of entropy concept that the power exponent r facilitates the introduction of entropy based constraints into fuzzy clustering algorithms. As an attractive and judicious application, DI-FCM is integrated with the fuzzy subspace clustering (FSC) algorithm so that a novel subspace clustering algorithm called double indices induced fuzzy subspace clustering (DI-FSC) algorithm is proposed for high dimensional data. In DI-FSC, the commonly-used Euclidean distance is replaced by the feature-weighted distance, which results in two fuzzy matrices in the objective function. Meanwhile, the convergence property of DI-FSC is also investigated. Experiments on the artificial data as well as the real text data were conducted and the experimental results show the effectiveness of the proposed algorithm. Jun Wang 0024, Shitong Wang 0001, Zhaohong Deng, Korris Fu-Lai Chung |
FUZZ-IEEE | 3 |
| 2012 | Fast Graph-Based Relaxed Clustering for Large Data Sets Using Minimal Enclosing BallabstractAlthough graph-based relaxed clustering (GRC) is one of the spectral clustering algorithms with straightforwardness and self-adaptability, it is sensitive to the parameters of the adopted similarity measure and also has high time complexity O(N(3)) which severely weakens its usefulness for large data sets. In order to overcome these shortcomings, after introducing certain constraints for GRC, an enhanced version of GRC [constrained GRC (CGRC)] is proposed to increase the robustness of GRC to the parameters of the adopted similarity measure, and accordingly, a novel algorithm called fast GRC (FGRC) based on CGRC is developed in this paper by using the core-set-based minimal enclosing ball approximation. A distinctive advantage of FGRC is that its asymptotic time complexity is linear with the data set size N. At the same time, FGRC also inherits the straightforwardness and self-adaptability from GRC, making the proposed FGRC a fast and effective clustering algorithm for large data sets. The advantages of FGRC are validated by various benchmarking and real data sets. Pengjiang Qian, Korris Fu-Lai Chung, Shitong Wang 0001, Zhaohong Deng |
IEEE Trans. Syst. Man Cybern. Part B | 4 |
| 2011 | Scalable TSK Fuzzy Modeling for Very Large Datasets Using Minimal-Enclosing-Ball ApproximationabstractIn order to overcome the difficulty in Takagi-Sugeno-Kang (TSK) fuzzy modeling for large datasets, scalable TSK (STSK) fuzzy-model training is investigated in this study based on the core-set-based minimal-enclosing-ball (MEB) approximation technique. The specified L2-norm penalty-based -insensitive criterion is first proposed for TSK-model training, and it is found that such TSK fuzzy-model training can be equivalently expressed as a center-constrained MEB problem. With this finding, an STSK fuzzy-model-training algorithm, which is called STSK, for large or very large datasets is then proposed by using the core-set-based MEB-approximation technique. The proposed algorithm has two distinctive advantages over classical TSK fuzzy-model training algorithms: The maximum space complexity for training is not reliant on the size of the training dataset, and the maximum time complexity for training is linear with the size of the training dataset, as confirmed by extensive experiments on both synthetic and real-world regression datasets. Zhaohong Deng, Kup-Sze Choi, Korris Fu-Lai Chung, Shitong Wang 0001 |
IEEE Trans. Fuzzy Syst. | 1 |
| 2010 | Enhanced soft subspace clustering integrating within-cluster and between-cluster information
Zhaohong Deng, Kup-Sze Choi, Korris Fu-Lai Chung, Shitong Wang 0001 |
Pattern Recognit. | 1 |
| 2010 | Robust Relief-Feature Weighting, Margin Maximization, and Fuzzy OptimizationabstractA latest advance in Relief-feature-weighting techniques is that the iterative procedure of Relief can be approximately expressed as a margin maximization problem, and therefore, its distinctive properties can be investigated with the help of optimization theory. Being motivated by this advance, the Relief-feature-weighting algorithm is investigated for the first time within a fuzzy-optimization framework. A new margin-based objective function that incorporates three fuzzy concepts, namely, fuzzy-difference measure, fuzzy-feature weighting, and fuzzy-instance force coefficient, is introduced. By the application of fuzzy optimization to this new margin-based objective function, several useful theoretical results are derived, based upon which, a set of robust Relief-feature-weighting algorithms are proposed for two-class data, multiclass data, and, then, online data. As demonstrated by extensive experiments in synthetic datasets, the University of California at Irvine (UCI)-benchmark datasets, cancer-gene-expression datasets, and face-image datasets, the proposed algorithms were found to be competitive with the state-of-the-art algorithms and robust for datasets with noise and/or outliers. Zhaohong Deng, Korris Fu-Lai Chung, Shitong Wang 0001 |
IEEE Trans. Fuzzy Syst. | 1 |
| 2009 | From Minimum Enclosing Ball to Fast Fuzzy Inference System Training on Large DatasetsabstractWhile fuzzy inference systems (FISs) have been extensively studied in the past decades, the minimum enclosing ball (MEB) problem was recently introduced to develop fast and scalable methods in pattern classification and machine learning. In this paper, the relationship between these two apparently different data modeling techniques is explored. First, based on the reduced-set density estimator, a bridge between the MEB problem and the FIS is established. Then, an important finding that the Mamdani-Larsen FIS (ML-FIS) can be translated into a special kernelized MEB problem, i.e., a center-constrained MEB problem under some conditions, is revealed. Thus, fast kernelized MEB approximation algorithms can be adopted to construct ML-FIS in an efficient manner. Here, we propose the use of a core vector machine (CVM), which is a fast kernelized MEB approximation algorithm for support vector machine (SVM) training, to accomplish this task. The proposed fast ML-FIS training algorithm has the following merits: (1) the number of fuzzy rules can be automatically determined by the CVM training and (2) fast ML-FIS training on large datasets can be achieved as the upper bound on the time complexity of learning the parameters in ML-FIS is linear with the dataset sizeNand the upper bound on the corresponding space complexity is theoretically independent ofN. Our experiments on simulated and real datasets confirm these advantages of the proposed training method, and demonstrate its superior robustness as well. This paper not only represents a very first study of the relationship between MEB and FIS, but it also points out the mutual transformation between kernel methods and FISs under the framework of the Gaussian mixture model and MEB. Korris Fu-Lai Chung, Zhaohong Deng, Shitong Wang 0001 |
IEEE Trans. Fuzzy Syst. | 2 |
| 2009 | An Adaptive Fuzzy-Inference-Rule-Based Flexible Model for Automatic Elastic Image RegistrationabstractIn this study, a fuzzy-inference-rule-based flexible model (FIR-FM) for automatic elastic image registration is proposed. First, according to the characteristics of elastic image registration, an FIR-FM is proposed to model the complex geometric transformation and feature variation in elastic image registration. Then, by introducing the concept of motion estimation and the corresponding sum-of-squared-difference (SSD) objective function, the parameter learning rules of the proposed model are derived for general image registration. Based on the likelihood objective function, particular attention is also paid to the derivation of parameter learning rules for the case of partial image registration. Thus, an FIR-FM-based automatic elastic image registration algorithm is presented here. It is distinguished by its 1) strong ability in approximating complex nonlinear transformation inherited from fuzzy inference; 2) efficiency and adaptability in obtaining precise model parameters through effective parameter learning rules; and 3) completely automatic registration process that avoids the requirement of manual control, as in many traditional landmark-based algorithms. Our experiments show that the proposed method has an obvious advantage in speed and is comparable in registration accuracy as compared with a state-of-the-art algorithm. Korris Fu-Lai Chung, Zhaohong Deng, Shitong Wang 0001 |
IEEE Trans. Fuzzy Syst. | 2 |
| 2008 | FRSDE: Fast reduced set density estimator using minimal enclosing ball approximation
Zhaohong Deng, Korris Fu-Lai Chung, Shitong Wang 0001 |
Pattern Recognit. | 1 |
| 2007 | A New Minimax Probability Based Classifier Using Fuzzy Hyper-EllipsoidabstractIn this paper, a new classifier called minimax-probability based fuzzy hyper-ellipsoid machine (MP-FHM) is proposed. It offers an alternative implementation of the minimax probability based classification with hyper plane and can be taken as an extended version of the ball-model based classifier. By the theorem proposed by Marshall and Qlkin, the training procedure of MP-FHM can be transformed into solving the corresponding unconstrained optimization problems, and thereby various optimization techniques can easily be adopted to solve them. In addition, the MP-FHM can be kernelized, and therefore it has strong nonlinear classification capabilities like other kernel-based classifiers. Various experiments were conducted and the results demonstrate that the proposed classifier is competitive with the state-of-the-art classifiers and is a very promising classification method. Zhaohong Deng, Korris Fu-Lai Chung, Shitong Wang 0001 |
IJCNN | 1 |
| 2006 | Clustering Analysis of Gene Expression Data based on Semi-supervised Visual Clustering Algorithm
Korris Fu-Lai Chung, Shitong Wang 0001, Zhaohong Deng, Chen Shu, Dewen Hu |
Soft Comput. | 3 |
| 2006 | Robust maximum entropy clustering algorithm with its labeling for outliers
Shitong Wang 0001, Korris Fu-Lai Chung, Zhaohong Deng, Dewen Hu, Xisheng Wu |
Soft Comput. | 3 |
| 2006 | CATSMLP: Toward a Robust and Interpretable Multilayer Perceptron With Sigmoid Activation FunctionsabstractEnhancing the robustness and interpretability of a multilayer perceptron (MLP) with a sigmoid activation function is a challenging topic. As a particular MLP, additive TS-type MLP (ATSMLP) can be interpreted based on single-stage fuzzy IF-THEN rules, but its robustness will be degraded with the increase in the number of intermediate layers. This paper presents a new MLP model called cascaded ATSMLP (CATSMLP), where the ATSMLPs are organized in a cascaded way. The proposed CATSMLP is a universal approximator and is also proven to be functionally equivalent to a fuzzy inference system based on syllogistic fuzzy reasoning. Therefore, the CATSMLP may be interpreted based on syllogistic fuzzy reasoning in a theoretical sense. Meanwhile, due to the fact that syllogistic fuzzy reasoning has distinctive advantage over single-stage IF-THEN fuzzy reasoning in robustness, this paper proves in an indirect way that the CATSMLP is more robust than the ATSMLP in an upper-bound sense. Several experiments were conducted to confirm such a claim. Korris Fu-Lai Chung, Shitong Wang 0001, Zhaohong Deng, Dewen Hu |
IEEE Trans. Syst. Man Cybern. Part B | 3 |