VLDB 2026 Research / reviewers in the wild / expert
Xingchen Hu 0001
dblp:170/3965
· DBLP profile ↗
42ranked-venue papers
15as first author
32since 2021 · last 2026
0000-0001-6879-5266ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 29 · 14 first-author · 19 since 2021Databases, data management, data science and information retrieval · 8 · 1 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DLME: A distillation mechanism from language models for knowledge graph embedding
Yuehang Si, Xingchen Hu 0001, Qing Cheng 0004, Jincai Huang 0001 |
Neurocomputing | 2 |
| 2026 | Few-Shot Specific Emitter Identification for Satellite IoE: A Novel Ground-Satellite Generative Learning FrameworkabstractAs a part of Internet of Everything (IoE), satellite communication networks have the problem of the low emitter identification performance due to limited signal samples. To address this issue, this article develops a novel ground-satellite generative learning (GSGL) framework that deploys the SEI model trained at the ground station to the low earth orbit (LEO) satellite for identification. In particular, we exploit the embedded temporal information losslessly by employing gramian angular field (GAF) to preserve the original features of signals, thereby highlighting radio frequency fingerprint (RFF) discriminability among devices. To overcome the limitation of insufficient samples, a hybrid discriminative mechanism driven generative learning is proposed to effectively generate high-fidelity GAF representations, and then the original and generated GAF representations are fed into a residual structure network for extracting RFFs. Considering channel states and generated deviation, we introduce a lightweight combination attention mechanism to reinforce the intra-class cohesion of RFFs by refining the quality of features. Simulation results demonstrate that the proposed framework can enhance inter-class separation and intra-class cohesion of RFFs to obtain excellent identification performance under the scenarios of insufficient samples. Furthermore, the proposed framework can achieve 31.56% accuracy improvement over state-of-the-art methods at 5 signal samples per category. Zhenhan Zhao, Jian Chen 0002, Yuchen Zhou 0001, Bingtao He, Min Hui, Xingchen Hu 0001, Huaiyu Tang |
IEEE Internet Things J. | 6 |
| 2026 | T-CT2CRP-SM: The exploration of dockless bike-sharing system rebalancing problems based on multi-modal signals in social networksabstractWith the continuous advancement of data processing technologies, decision-making problems based on multi-modal signal systems (MMSSs) can be effectively addressed. In this context, rebalancing problems in dockless bike-sharing systems (DBSSs) within social networks increasingly rely on MMSSs. However, MMSSs introduce the issue of declining information quality, making it essential to ensure the reliability of both signals and models. Specifically, the paper addresses these challenges by exploring methods to reduce the impact of low-quality signals in MMSSs, leading to the design of a trustworthy multi-modal signal processing (TMSP) model. Yet, two main challenges remain in building the model, i.e., the lack of suitable frameworks to accurately characterize MMSSs, and insufficient coupling between clustering analysis and the consensus-reaching process (CRP). To address these challenges, the signal reliability is first improved by processing MMSSs with complex intuitionistic fuzzy sets (CIFSs). Then, based on the MMSS, a trustworthy clustering and two-stage, two-index CRP model based on similarity measurement (T-CT2CRP-SM) is constructed. Subsequently, experiments select a real-world MMSS-based DBSS rebalancing problem as the scenario and illustrate detailed decision-making steps. Finally, multiple experimental analyses are conducted to further validate the reliability and stability of the constructed model. Chao Zhang 0046, Anna Wang 0003, Wentao Li 0004, Xingchen Hu 0001 |
Signal Process. | 5 |
| 2026 | Structure identification of missing data: a perspective from granular computing
Yinghua Shen, Xingchen Hu 0001, Witold Pedrycz, Zhi Xiao |
Soft Comput. | 3 |
| 2026 | Zero-Shot Event Causality Identification via Multisource Evidence Fuzzy Aggregation With Large Language ModelsabstractEvent causality identification (ECI) aims to detect causal relationships between events in textual contexts. Existing ECI models predominantly rely on supervised methodologies, suffering from dependence on large-scale annotated data. Although large language models (LLMs) enable zero-shot ECI, they are prone to causal hallucination—erroneously establishing spurious causal links. To address these challenges, we propose MEFA, a novel zero-shot ECI model based on multisource evidence fuzzy aggregation. First, we decompose causality reasoning into three main tasks (temporality determination, necessity analysis, and sufficiency verification) complemented by three auxiliary tasks. Second, leveraging meticulously designed prompts, we guide LLMs to generate uncertain responses and deterministic outputs. Finally, we quantify LLM's responses of subtasks and employ fuzzy aggregation to integrate these evidence for causality scoring and causality determination. Extensive experiments on three benchmarks demonstrate that MEFA outperforms second-best unsupervised baselines by 6.2% in$F1$-score and 9.3% in precision, while significantly reducing hallucination-induced errors. In-depth analysis verify the effectiveness of task decomposition and the superiority of fuzzy aggregation. Zefan Zeng, Qing Cheng 0004, Xingchen Hu 0001, Wentao Li 0004, Weiping Ding 0001, Zhong Liu 0002 |
IEEE Trans. Fuzzy Syst. | 3 |
| 2026 | Contrastive and Dual Adversarial Representation Learning for Multi-View ClusteringabstractMulti-View Clustering (MVC) has gained increasing attention due to its ability to effectively leverage the complementary information of multi-view data. Despite the success of existing MVC methods in many real-world applications, they often overlook the discrepancy of view-specific latent distribution and struggle to ensure the completeness of the multi-view data. To address these challenges and harness the powerful feature extraction capability of deep networks, we propose a novel Contrastive and Dual Adversarial Representation Learning method for Multi-view Clustering, termed as CDARL, to solve multi-view clustering problems with both complete and incomplete multi-view data. Specifically, CDARL employs alternating adversarial and contrastive learning to align the view-specific representations, driving them into the same semantic latent space to minimize the discrepancy in view-specific distributions. In addition, a consensus latent representation is learned by an adaptive fusion block that integrates information from multiple views. The consensus representation is further refined through adversarial learning modeling the transformation of the standard Gaussian distribution to the original data distribution. Moreover, the proposed method incorporates an imputation strategy designed to handle the incomplete multi-view data clustering task. This strategy utilizes both reconstructed samples and cross-view neighbors to impute missing views from the latent space and the original space, thereby preserving clustering information, which ensures the quality and feasibility of the imputed samples. Experimental results on six widely used datasets have verified the competitiveness of the proposed CDARL method against state-of-the-art methods in MVC problems with complete and incomplete multi-view data. Code is available athttps://github.com/xywy220/CDARL-MVC. Yanwanyu Xi, Chang Tang, Junjie Huang 0001, Xingchen Hu 0001, Yuanyuan Liu 0004, Xinwang Liu 0002 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2026 | FedMPS: Federated Learning in a Synergy of Multi-Level Prototype-Based Contrastive Learning and Soft Label GenerationabstractFederated learning (FL) facilitates collaborative training among multiple clients while preserving data privacy by eliminating raw data transmission. However, the inherent data heterogeneity among participants induces bias during collaborative learning, significantly degrading the performance of local models. Existing FL solutions face critical challenges in achieving efficient knowledge transmission, particularly with respect to insufficient information extraction or excessive communication costs, which result in slow convergence and inferior performance. To address these limitations, we propose a novel FL framework in a synergy of multi-level prototype-based contrastive learning (CL) and soft label generation, named FedMPS. The proposed method first constructs multi-level prototypes from different layers of the model to capture semantic information in high-level features and detailed information in low-level features. These prototypes are then utilized through CL to enhance intra-class discriminability and intra-class consistency in the feature space. In addition, a prototype-guided soft label generation module is introduced to model latent interclass relationships in the output space. Instead of exchanging model parameters, FedMPS transmits only prototypes and soft labels, effectively reducing global knowledge shift and communication costs. Extensive experimental studies on six publicly available datasets validate the effectiveness of the proposed method when compared to the current state-of-the-art FL approaches. The code is available at github.com/wenxinyang1026/FedMPS. Wenxin Yang, Xingchen Hu 0001, Xiubin Zhu, Rouwan Wu, Witold Pedrycz, Xinwang Liu 0002, Jincai Huang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2025 | EASEMVC: Efficient Dual Selection Mechanism for Deep Multi-View ClusteringabstractMulti-view clustering (MVC) has emerged as a leading paradigm in unsupervised learning, gaining significant attention. Central to this framework is the concept of view-pair contrastive learning, which aims to maximize mutual information between pairs of views, thereby facilitating consistent latent representations. Nevertheless, two critical challenges remain: i) Identifying the most suitable pairs of views for contrastive learning becomes difficult when more than two views are available, especially in the absence of prior knowledge; ii)Including all available views in contrastive learning can degrade performance due to the presence of low-quality views. To address these issues, we propose a novel mechanism, EASEMVC(Efficient DuAl Selection MEchanism for Deep Multi-View Clustering). EASEMVC begins by constructing a view graph using the Optimal Transport (OT) distance between bipartite graphs of individual views. A view selection module is then designed to perform efficient view-level selection based on the topological relationships within the view graph. Additionally, a cross-view sample graph is built at the sample level, where the topological relationships among samples are used to generate reliable learning weights. Leveraging the selected view pairs and sample weights, contrastive learning is employed to obtain consistent representations across views. Extensive experiments across six benchmark datasets demonstrate that EASEMVC outperforms current state-of-the-art methods. Baili Xiao, Zhibin Dong, Ke Liang 0006, Suyuan Liu, Siwei Wang 0001, Tianrui Liu 0001, Xingchen Hu 0001, En Zhu, Xinwang Liu 0002 |
CVPR | 7 |
| 2025 | Measuring the Impact of Rotation Equivariance on Aerial Object DetectionabstractDue to the arbitrary orientation of objects in aerial images, rotation equivariance is a critical property for aerial object detectors. However, recent studies on rotation-equivariant aerial object detection remain scarce. Most detectors rely on data augmentation to enable models to learn approximately rotation-equivariant features. A few detectors have constructed rotation-equivariant networks, but due to the breaking of strict rotation equivariance by typical downsampling processes, these networks only achieve approximately rotation-equivariant backbones. Whether strict rotation equivariance is necessary for aerial image object detection remains an open question. In this paper, we implement a strictly rotation-equivariant backbone and neck network with a more advanced network structure and compare it with approximately rotation-equivariant networks to quantitatively measure the impact of rotation equivariance on the performance of aerial image detectors. Additionally, leveraging the inherently grouped nature of rotation-equivariant features, we propose a multi-branch head network that reduces the parameter count while improving detection accuracy. Based on the aforementioned improvements, this study proposes the Multi-branch head rotation-equivariant single-stage Detector (MessDet), which achieves state-of-the-art performance on the challenging aerial image datasets DOTA-v1.0, DOTA-v1.5 and DIOR-R with an exceptionally low parameter count. Xiuyu Wu, Xiubin Zhu, Lan Yang 0007, Jiyuan Liu 0003, Xingchen Hu 0001 |
ICCV | 6 |
| 2025 | LRGR: Self-Supervised Incomplete Multi-View Clustering via Local Refinement and Global RealignmentabstractIncomplete Multi-View Clustering (IMVC) aims to explore comprehensive representations from multiple views with missing samples. Recent studies have revealed that IMVC methods benefit from Graph Convolutional Network (GCN) in achieving robust feature imputation and effective representation learning. Despite these notable improvements, GCN imputation methods often cause a distribution shift between the imputed and original representations, particularly when the neighbors of the imputed nodes are assigned to different groups. Moreover, GCN learning methods tend to produce homogeneous imputed representations, which blur cluster boundaries and hinder effective discriminative clustering. To remedy these challenges, the Local Refinement and Global Realignment (LRGR) Self-supervised model is proposed for incomplete multi-view clustering, which includes two stages. In the first stage, a local imputed refinement module is designed to enhance the versatility of imputed representations through cross-view contrastive learning guided by view-specific prototypes. In the second stage, a global realignment module is introduced to achieve semantic consistency across views, alleviating distribution shifts by leveraging pseudo-labels and their corresponding confidence scores as guidance. Experiments on five widely used multi-view datasets demonstrate the competitiveness and superiority of our method compared to state-of-the-art approaches. Yanwanyu Xi, Chang Tang, Xingchen Hu 0001, Yuanyuan Liu 0004, Xinwang Liu 0002 |
IJCAI | 4 |
| 2025 | Federated Incomplete Multi-view Clustering with Individual Structure Preservation and Central Representation Tensorization
Yan Li 0003, Xingchen Hu 0001, Jiyuan Liu 0003, Zhong Liu 0002 |
ACM Multimedia | 2 |
| 2025 | SALVG: Latent Variable Gene Augmented Graph Learning for Multi-View Clustering in Spatial TranscriptomicsabstractSpatial transcriptomics technologies enable the integration of gene expression profiles with spatial context, facilitating a deeper understanding of tissue architecture through downstream tasks such as clustering. However, existing approaches predominantly focus on highly variable genes (HVGs), while the informative structural and contextual signals embedded in low variability genes (LVGs) remain largely underutilized. To bridge this gap, we propose SALVG (Spatial Augmentation via Latent Variable Genes), a novel and plug-and-play framework that leverages LVG-derived structural priors to enhance HVG representation learning for spatial clustering. Specifically, SALVG constructs spatial, feature, and combined graphs for both HVGs and LVGs, and introduces two graph-based augmentation strategies to inject LVG information into HVG graphs. The first strategy enhances the HVG combined graph directly using the LVG combined graph, while the other individually augments HVG spatial and feature graphs with their LVG counterparts before fusing them into a new combined representation. These enhanced graph structures are subsequently employed for downstream clustering. To the best of our knowledge, SALVG is the first framework to exploit LVG signals for assisting HVG-centric spatial transcriptomics clustering, effectively capturing complementary structural and contextual cues. Experiments on multiple benchmarks demonstrate its effectiveness, robustness, and transferability. Case studies further confirm that LVG-derived structure enhances biological interpretability by revealing coherent spatial and cellular patterns. Ke Liang 0006, Lingyuan Meng, Xingchen Hu 0001, Xinwang Liu 0002, Wanwei Liu, Kunlun He |
ACM Multimedia | 4 |
| 2025 | SparseMVC: Probing Cross-view Sparsity Variations for Multi-view ClusteringabstractExisting multi-view clustering methods employ various strategies to address data-level sparsity and view-level dynamic fusion. However, we identify a critical yet overlooked issue: varying sparsity across views. Cross-view sparsity variations lead to encoding discrepancies, heightening sample-level semantic heterogeneity and making view-level dynamic weighting inappropriate. To tackle these challenges, we propose Adaptive Sparse Autoencoders for Multi-View Clustering (SparseMVC), a framework with three key modules. Initially, the sparse autoencoder probes the sparsity of each view and adaptively adjusts encoding formats via an entropy-matching loss term, mitigating cross-view inconsistencies. Subsequently, the correlation-informed sample reweighting module employs attention mechanisms to assign weights by capturing correlations between early-fused global and view-specific features, reducing encoding discrepancies and balancing contributions. Furthermore, the cross-view distribution alignment module aligns feature distributions during the late fusion stage, accommodating datasets with an arbitrary number of views. Extensive experiments demonstrate that SparseMVC achieves state-of-the-art clustering performance. Our framework advances the field by extending sparsity handling from the data-level to view-level and mitigating the adverse effects of encoding discrepancies through sample-level dynamic weighting. The source code is publicly available at https://github.com/cleste-pome/SparseMVC. Ruimeng Liu, Xin Zou 0001, Chang Tang, Xingchen Hu 0001, Kun Sun 0002, Xinwang Liu 0002 |
NeurIPS | 5 |
| 2025 | Coherence mode: Characterizing local graph structural information for temporal knowledge graph
Yuehang Si, Xingchen Hu 0001, Qing Cheng 0004, Xinwang Liu 0002, Jincai Huang 0001 |
Inf. Sci. | 2 |
| 2025 | KoSEL: Knowledge subgraph enhanced large language model for medical question answering
Zefan Zeng, Qing Cheng 0004, Xingchen Hu 0001, Xinwang Liu 0002, Kunlun He, Zhong Liu 0002 |
Knowl. Based Syst. | 3 |
| 2025 | Deep Temporal Graph Clustering: A Comprehensive Benchmark and DatasetsabstractTemporal Graph Clustering (TGC) is a new task with little attention, focusing on node clustering in temporal graphs. Compared with existing static graph clustering, it can find the balance between time requirement and space requirement (Time-Space Balance) through the interaction sequence-based batch-processing pattern. However, there are two major challenges that hinder the development of TGC, i.e., inapplicable clustering techniques and inapplicable datasets. To address these challenges, we propose a comprehensive benchmark, called BenchTGC. Specially, we design a BenchTGC Framework to illustrate the paradigm of temporal graph clustering and improve existing clustering techniques to fit temporal graphs. In addition, we also discuss problems with public temporal graph datasets and develop multiple datasets suitable for TGC task, called BenchTGC Datasets. According to extensive experiments, we not only verify the advantages of BenchTGC, but also demonstrate the necessity and importance of TGC task. We wish to point out that the dynamically changing and complex scenarios in real world are the foundation of temporal graph clustering. Meng Liu 0014, Ke Liang 0006, Siwei Wang 0001, Xingchen Hu 0001, Sihang Zhou 0001, Xinwang Liu 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | A Vertical Federated Multiview Fuzzy Clustering Method for Incomplete DataabstractMulti-view fuzzy clustering (MVFC) has gained widespread adoption owing to its inherent flexibility in handling ambiguous data. The proliferation of privatization devices has driven the emergence of new challenge in MVFC researches. Federated learning, a technique that can jointly train without directly using raw data, has gain significant attention in decentralized MVFC. However, their applicability depends on the assumptions of data integrity and independence between different views. In fact, while within distributed environments, data typically exhibits two challenging problems: (1) multiple views within a single client; (2) incomplete data. Existing methods exhibit limitations in effectively addressing these challenges. Hence, in this study, we aim at achieving the effective clustering for incomplete data by a novel vertical federated MVFC framework. Specifically, a unified clustering framework is designed to capture both local client learning and global server training. For the local client learning, the data reconstruction strategy and prototype alignment strategy are introduced to ensure the preservation of data structure and refinement of clustering relationships, which mitigates the impact of incomplete data. Meanwhile, the global training process implements aggregation based on client-specific information. The whole process is realized based on the unified fuzzy clustering framework, promoting collaborative learning between client-specific and server information. Theoretical analyses and extensive experiments are carefully conducted to validate the effectiveness and efficiency of the proposed method from multiple perspectives. Xingchen Hu 0001, Shengju Yu, Weiping Ding 0001, Witold Pedrycz, Chai Kiat Yeo, Zhong Liu 0002 |
IEEE Trans. Fuzzy Syst. | 2 |
| 2025 | Dynamic Ensemble Framework for Imbalanced Data ClassificationabstractDynamic ensemble has significantly greater potential space to improve the classification of imbalanced data compared to static ensemble. However, dynamic ensemble schemes are far less successful than static ensemble methods in the imbalanced learning field. Through an in-depth analysis on the behavior characteristics of dynamic ensemble, we find that there are some important problems that need to be addressed to release the full potential of dynamic ensemble, including but not limited to, correcting the component classifiers’ bias towards the majority classes, increasing the proportions of the positive classifiers (i.e., the component classifiers making correct prediction) for difficult samples, and providing the accurate competence estimations on the hard-to-classify samples w.r.t the classifier pool. Inspired by these, we propose a Dynamic Ensemble Framework for imbalanced data classification (imDEF). imDEF first uses the data generation method OREM$\mathrm{_{G}}$to generate multiple artificial synthetic datasets, which have diverse class distributions by rebalancing the original imbalanced data. Based on each of such synthetic datasets, imDEF then utilizes a Classification Error-aware Self-Paced Sampling Ensemble (SPSE$\mathrm{_{CE}}$) method to gradually focus more on difficult samples, to create a low-biased classifier pool and increase the proportions of the positive classifiers for the difficult samples. Finally, imDEF constructs a referee system to achieve the competence estimations by leveraging an Ensemble Margin-aware Self-Paced Sampling Ensemble (SPSE$\mathrm{_{EM}}$) method. SPSE$\mathrm{_{EM}}$incrementally strengthens the learning of the hard-to-classify samples, so that the competent levels of component classifiers could be estimated accurately. Extensive experiments demonstrate the effectiveness of imDEF. The source codes have been made publicly available on GitHub. Tuanfei Zhu, Xingchen Hu 0001, Xinwang Liu 0002, En Zhu, Xinzhong Zhu |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2024 | Aligning the Representation of Knowledge Graph and Large Language Model for Causal Question AnsweringabstractCausal Question Answering (CQA) is essential for knowledge discovery, focusing on the intricate dynamics between events and entities without predefined contexts. Despite advancements of CQA models through Knowledge Graphs (KGs) and Pre-Trained Language Models (PLMs), existing approaches are hindered by knowledge conflict, insufficient capacity, and limitations in information fusion. Large Language Models (LLMs) have significantly improved natural language understanding and reasoning but often suffer from causal hallucinations. To address these challenges, we introduce KLop, a framework that aligns representations of Causal Knowledge Graph (CKG) and Large Language Models for CQA. KLop pre-trains a graph embedding model for entity embedding and uses a frozen LLM for text embedding. The main components of KLop are the descriptor module and the aligner module. The descriptor leverages descriptive texts generated by LLMs to create training data for knowledge alignment, while the aligner utilizes self-attention to train query tokens for modality alignment. Experiments on public CQA datasets validate that KLop outperforms various advanced baselines in reasoning accuracy, as well as achieving causal knowledge integration and joint reasoning. Zefan Zeng, Qing Cheng 0004, Xingchen Hu 0001, Zhong Liu 0002, Jingke Shen, Yahao Zhang |
IEEE Big Data | 3 |
| 2024 | Reliable Attribute-missing Multi-view Clustering with Instance-level and feature-level Cooperative ImputationabstractMulti-view clustering (MVC) constitutes a distinct approach to data mining within the field of machine learning. Due to limitations in the data collection process, missing attributes are frequently encountered. However, existing MVC methods primarily focus on missing instances, showing limited attention to missing attributes. A small number of studies employ the reconstruction of missing instances to address missing attributes, potentially overlooking the synergistic effects between the instance and feature spaces, which could lead to distorted imputation outcomes. Furthermore, current methods uniformly treat all missing attributes as zero values, thus failing to differentiate between real and technical zeroes, potentially resulting in data over-imputation. To mitigate these challenges, we introduce a novel Reliable Attribute-Missing Multi-View Clustering method (RAM-MVC). Specifically, feature reconstruction is utilized to address missing attributes, while similarity graphs are simultaneously constructed within the instance and feature spaces. By leveraging structural information from both spaces, RAM-MVC learns a high-quality feature reconstruction matrix during the joint optimization process. Additionally, we introduce a reliable imputation guidance module that distinguishes between real and technical attribute-missing events, enabling discriminative imputation. The proposed RAM-MVC method outperforms nine baseline methods, as evidenced by real-world experiments using single-cell multi-view data. Dayu Hu, Suyuan Liu, Jun Wang 0118, Junpu Zhang, Siwei Wang 0001, Xingchen Hu 0001, Xinzhong Zhu, Chang Tang, Xinwang Liu 0002 |
ACM Multimedia | 6 |
| 2024 | Discriminative embedded multi-view fuzzy C-means clustering for feature-redundant and incomplete data
Yan Li 0003, Xingchen Hu 0001, Tuanfei Zhu, Jiyuan Liu 0003, Xinwang Liu 0002, Zhong Liu 0002 |
Inf. Sci. | 2 |
| 2024 | An Efficient Federated Multiview Fuzzy C-Means Clustering MethodabstractMulti-view clustering has been received considerable attention due to the widespread collection of multi-view data from diverse domains and sources. However, storing multi-view data across multiple devices in many real scenarios poses significant challenges for efficient data analysis. Federated Learning framework enables collaborative machine learning on distributed devices while preserving privacy constraints. Even though there have been intensive algorithms on multi-view fuzzy clustering, federated multi-view fuzzy clustering has not been adequately investigated so far. In this study, we first develop the federated learning mode into multi-view fuzzy clustering and realize the federated optimization procedure, called Federated Multiview Fuzzy C-Means clustering (FedMVFCM). Then, we design an original strategy of consensus prototype learning during federated multi-view fuzzy clustering. It is termed as Federated Multi-view Fuzzy c-means consensus Prototypes Clustering (FedMVFPC). We also further develop the federated alternative optimization algorithm with proven convergence. This study also introduces the notion of clustering prototype communication within the federated learning framework, and integrates the clustering prototypes of different views into a unified optimization formulation. The experimental studies on various benchmark datasets demonstrate that the proposed FedMVFPC method improves the federated clustering performance and efficiency. It achieves comparable or better clustering performance against the existing state-of-the-art multi-view clustering algorithms Xingchen Hu 0001, Jindong Qin, Yinghua Shen, Witold Pedrycz, Xinwang Liu 0002, Jiyuan Liu 0003 |
IEEE Trans. Fuzzy Syst. | 1 |
| 2024 | A Design of Fuzzy Rule-Based Classifier for Multiclass Classification and Its Realization in Horizontal Federated LearningabstractPattern recognition plays an important role in the process of knowledge discovery. The construction of easily describable and interpretable classification rules is of vital importance in pattern recognition. In this study, we propose a development of fuzzy rule-based classifier for multiclass classification problems and elaborate on a privacy-preserving realization of the proposed methodology in the presence of decentralized datasets. Fuzzy rule-based models provide an effective and efficient alternative for characterizing the complex relationship between the input variables and target classes. An overall design process of the proposed classifier consists of two main phases: (a) formation of information granules (clusters) to reveal the underlying structure of the training data, and (b) construction of local classification rules whose outputs reflect the probability distribution of the input data over all the classes. The constructed information granules form a backbone of the architecture of the classifier while the optimization of the parameters of local rules is carried out through using a gradient descent method with the guidance of the cross-entropy loss function. Furthermore, a federated gradient-based optimization mechanism is utilized to construct fuzzy classifier in a privacy-preserving approach. The originalities of the proposed methodology are twofold: first, a design of fuzzy classifier through the synergy of cluster-centric architecture and the cross-entropy loss function is presented. Second, we augment the proposed fuzzy classifier based on the concept of federated learning such that it can learn from distributed data without sacrificing data security and confidentiality. Experiments are carried out on a two-dimensional synthetic dataset and a number of real-world datasets. Experimental results show the excellent classification capability of the proposed classifier realized in the centralized way and in the federated learning environment. Xingchen Hu 0001, Xiubin Zhu, Lan Yang 0007, Witold Pedrycz, Zhiwu Li 0001 |
IEEE Trans. Fuzzy Syst. | 1 |
| 2024 | A Granular Aggregation of Multifaceted Gaussian Process ModelsabstractThis study focuses on the construction of granular Gaussian process models completed at different levels of granularity and the emergence of higher-type granular outputs through aggregating the individual prediction results. Each Gaussian process model is instantiated utilizing granular data (or information granules) to enhance algorithmic efficiency and can be tailored to specific levels of precision (granularity). The overall design methodology emphasizes human centricity in system modeling by focusing on both the interpretability and accuracy of the resulting models. First, clustering algorithms are applied to construct information granules that provide a comprehensive overview of the experimental evidence. As the number of information granules grows, the existing knowledge imbedded within data could be perceived and described at increased levels of details. Information granules are built in an augmented feature space constructed by concatenating the input and output variables. Next, Gaussian process models are constructed on a basis of the information granules formed at different levels of abstraction. Subsequently, the confidence intervals are transformed to intervals and the reconciliation of the predictions produced by individual models, which offer different perspectives on the system, leads to the emergence of more abstract entities (such as type-2 intervals/fuzzy sets, etc.) rather than plain numbers. The efficacy of the comprehensive model is measured by the coverage and specificity criteria of the granular outputs. Experimental studies conducted on a synthetic dataset and a number of real-world datasets validated the effectiveness and adaptability of the proposed methodology. Lan Yang 0007, Xiubin Zhu, Witold Pedrycz, Zhiwu Li 0001, Xingchen Hu 0001 |
IEEE Trans. Fuzzy Syst. | 5 |
| 2024 | A Development of Fuzzy-Rule-Based Regression Models Through Using Decision TreesabstractThis article presents a design and realization of fuzzy rule-based regression models based on standard decision trees. A two-phase design of rule-based model is offered in this study to provide a good alternative to cope with high dimensional data. We first build a standard decision tree on the basis of variables in order to discover homogeneous subsets of the data. Subsequently, a collection of fuzzy rules is induced by the decision tree with the aim of reflecting the underlying phenomenon. The calculation of membership degrees and the refinement of fuzzy rules on the basis of data located in each partition exhibit a substantial level of originality and innovation. The introduction of fuzziness into decision rules helps to characterize and quantify the continuous change of output values near the boundary areas. The constructed fuzzy rules could efficiently handle the ambiguity and vagueness in the experimental evidence and offer an accurate characterization of the nonlinearities of the input–output relationships. The developed fuzzy models could achieve much higher prediction accuracy in comparison with traditional decision trees of the same size and fuzzy rule-based models with the same number of rules. Another advantage of the proposed methodology comes with the evident readability of the formed fuzzy rules. A series of experiments is reported to demonstrate the superiority of the proposed architecture of fuzzy rule-based models over traditional fuzzy rule-based models and decision trees. Xiubin Zhu, Xingchen Hu 0001, Lan Yang 0007, Witold Pedrycz, Zhiwu Li 0001 |
IEEE Trans. Fuzzy Syst. | 2 |
| 2024 | Cooperated Truck-Drone Routing With Drone Energy Consumption and Time WindowsabstractConsidering customer time windows, multiple trucks, each equipped with a multi-visit drone, are employed. For realism, the drone energy consumption and the impact of payload variation on the energy consumption rate are considered. A Mixed Integer Linear Programming (MILP) model is developed to formulate the problem. A novel concept called “Segment” is introduced to promote the cooperation between trucks and drones. Based on this, a heuristic is designed where the drone and truck routes are constructed synchronously. The variable neighborhood search algorithm is integrated with simulated annealing to enhance solutions further. The effectiveness of the proposed algorithm is validated through both Solomon instances and practical cases. Jianmai Shi, Xingchen Hu 0001, Witold Pedrycz, Zhong Liu 0002 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2024 | Application of Gradient Boosting in the Design of Fuzzy Rule-Based Regression ModelsabstractThis study is devoted to the design of gradient boosted fuzzy rule-based models for regression problems. Fuzzy rule-based models are built on the basis of information granules formed in the input and output spaces whose structure involves a family of conditional ‘if-then’ statements. The architecture of fuzzy rule-based models contributes to the realization of a sound tradeoff between modeling accuracy and interpretability and computing overhead. Gradient boosting paradigm has emerged as a powerful learning method realized through sequentially fitting additive base learners to current residuals in the steepest descent way. However, surprisingly, studies on the design and analysis of gradient boosted fuzzy rule-based models are still lacking. In this study, fuzzy rule-based model is regarded as a base learner. Different loss functions and their influence on the performance of the final models are explored. We also thoroughly investigate an impact of the initial quality of the rule-based model (implied by the number of rules) on the process of gradient boosting. The performance of the proposed approach is illustrated by a series of experimental studies concerning synthetic and publicly available datasets. Xingchen Hu 0001, Xiubin Zhu, Xinwang Liu 0002, Witold Pedrycz |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2023 | Multi-View Fuzzy Classification With Subspace Clustering and Information GranulesabstractMulti-view learning becomes increasingly attractive and promising because multimodal or multi-view data are commonly encountered in real-world applications. In this study, we develop a novel multi-view Takagi–Sugeno–Kang (TSK) fuzzy system framework to handle classification problems for such data. We propose an anchor and graph subspace clustering strategy to discover and represent the actual latent data distribution for each view separately. In this way, the discriminate anchors (landmarks) are learned to capture the main structure of the multi-view data. This strategy also provides a computationally efficient clustering algorithm with respect to the number of instances. These resulting anchors are formed as the prototypes of information granules (IGs) for fuzzy modeling. Then we construct an information-granule-based multi-view TSK fuzzy classification model inherited from the natural interpretability of fuzzy rule-based systems. Concretely, the relationship between the multi-view input and label output spaces is depicted by IGs-oriented fuzzy rules. The experimental studies involve various commonly used benchmark datasets, which indicate that our proposed method achieves comparable or better performance compared to the state-of-the-art algorithms. Xingchen Hu 0001, Xinwang Liu 0002, Witold Pedrycz, Qing Liao 0001, Yinghua Shen, Yan Li 0003, Siwei Wang 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2022 | Granular Fuzzy Rule-Based Modeling With Incomplete Data RepresentationabstractIncomplete data are frequently encountered and bring difficulties when it comes to further processing. The concepts of granular computing (GrC) help deliver a higher level of abstraction to address this problem. Most of the existing data imputation and related modeling methods are of numeric nature and require prior numeric models to be provided. The underlying objective of this study is to introduce a novel and straightforward approach that uses information granules as a vehicle to effectively represent missing data and build granular fuzzy models directly from resulting hybrid granular and numeric data. The evaluation and optimization of this method are guided by the principle of justifiable granularity engaging the coverage and specificity criteria and carried out with the help of particle swarm optimization. We provide a collection of experimental studies using a synthetic dataset and several publicly available real-world datasets to demonstrate the feasibility and analyze the main features of this method. Xingchen Hu 0001, Yinghua Shen, Witold Pedrycz, Yan Li 0003, Guohua Wu 0001 |
IEEE Trans. Cybern. | 1 |
| 2022 | Identification of Fuzzy Rule-Based Models With Collaborative Fuzzy ClusteringabstractFuzzy rule-based models (FRBMs) are sound constructs to describe complex systems. However, in reality, we may encounter situations, where the user or owner of a system only owns either the input or output data of that system (the other part could be owned by another user); and due to the consideration of data privacy, he/she could not obtain all the needed data to build the FRBMs. Since this type of situation has not been fully realized (noticed) and studied before, our objective is to come up with some strategy to address this challenge to meet the specific privacy consideration during the modeling process. In this study, the concept and algorithm of the collaborative fuzzy clustering (CFC) are applied to the identification of FRBMs, describing either multiple-input-single-output (MISO) or multiple-input-multiple-output (MIMO) systems. The collaboration between input and output spaces based on their structural information (conveyed in terms of the corresponding partition matrices) makes it possible to build FRBMs when input and output data could not be collected and used in unison. Surprisingly, on top of this primary pursuit, with the collaboration mechanism the input and output spaces of a system are endowed with an innovative way to comprehensively share, exchange, and utilize the structural information between each other, which results in their more relevant structures that guarantee better model performance compared with performance produced by some state-of-the-art modeling strategies. The effectiveness of the proposed approach is demonstrated by experiments on a series of synthetic and publicly available datasets. Xingchen Hu 0001, Yinghua Shen, Witold Pedrycz, Xianmin Wang, Adam Gacek, Bingsheng Liu |
IEEE Trans. Cybern. | 1 |
| 2021 | Fuzzy Rule-Based Models: A Design with Prototype Relocation and Granular Generalization
Yan Li 0003, Chao Chen 0017, Xingchen Hu 0001, Jindong Qin |
Inf. Sci. | 3 |
| 2021 | Information granule-based classifier: A development of granular imputation of missing data
Xingchen Hu 0001, Witold Pedrycz, Keyu Wu 0004, Yinghua Shen |
Knowl. Based Syst. | 1 |
| 2019 | Allocation of Information Granularity: A Multi-Objective Evolutionary Optimization Using Conflict InformationabstractGranular Computing (GrC) and its related granular modeling methodologies have received much attention recently to take advantage of information granules. The principle of justifiable granularity fundamentally guides allocation of information granularity along with its optimization. In the existing literature, two conflict criteria coverage and specificity are optimized by aggregating them as one indicator. In this study, we thoughtfully select evolutionary multi-objective optimization (EMO) approaches for the allocation of information granularity and pave a way to design a new EMO framework considering an analysis of the conflict information of these two objectives. To demonstrate the usefulness of EMO and our proposed algorithms, we present a series of experimental studies based on a typical example of granular fuzzy rule-based models. The experimental results indicate that EMO is more efficient to find optimal solutions of allocation of information granularity. Moreover, the proposed EMO method is more feasible to find a set of solutions which exhibits superior coverage or specificity as possible. Xingchen Hu 0001, Huangke Chen, Chao Chen 0017, Boliang Sun, Jincai Huang 0001, Kuihua Huang |
FUZZ-IEEE | 1 |
| 2019 | Clustering of Information Granules in Hotspot IdentificationabstractConceptually and algorithmically, hotspots could be regarded as information granules. In this study, we propose an aggregation of Fuzzy C-Means (FCM) algorithm and the principle of justifiable granularity (PJG) as a new approach to forming hotspots. With the proposed method, the quality of the hotspots formed in this manner could also be provided as an additional information to the decision makers. Moreover, a weighted granular clustering method is presented to further abstract the constructed hotspots, and this delivers a higher level of abstraction of the phenomenon of interest. A collection of synthetic data is used to show the proposed process of identifying the hotspots, and to demonstrate its differences with some other representative hotspot identification methods. Besides, real-world data are also used to illustrate the performance of the proposed method. Yinghua Shen, Witold Pedrycz, Ronei Marcos de Moraes, Xingchen Hu 0001, Xianmin Wang, Adam Gacek |
FUZZ-IEEE | 4 |
| 2019 | Fuzzy rule-based models with randomized development mechanisms
Xingchen Hu 0001, Witold Pedrycz, Dianhui Wang 0001 |
Fuzzy Sets Syst. | 1 |
| 2019 | Random ensemble of fuzzy rule-based models
Xingchen Hu 0001, Witold Pedrycz, Xianmin Wang |
Knowl. Based Syst. | 1 |
| 2018 | Fuzzy classifiers with information granules in feature space and logic-based computing
Xingchen Hu 0001, Witold Pedrycz, Xianmin Wang |
Pattern Recognit. | 1 |
| 2017 | Fuzzy rule-based models with interactive rules and their granular generalization
Xingchen Hu 0001, Witold Pedrycz, Oscar Castillo 0001, Patricia Melin |
Fuzzy Sets Syst. | 1 |
| 2017 | From fuzzy rule-based models to their granular generalizations
Xingchen Hu 0001, Witold Pedrycz, Xianmin Wang |
Knowl. Based Syst. | 1 |
| 2017 | Development of granular models through the design of a granular output spaces
Xingchen Hu 0001, Witold Pedrycz, Xianmin Wang |
Knowl. Based Syst. | 1 |
| 2017 | Granular Fuzzy Rule-Based Models: A Study in a Comprehensive Evaluation and Construction of Fuzzy ModelsabstractFuzzy models are regarded as numeric constructs and as such are optimized and evaluated at the numeric level. In this study, we depart from this commonly accepted position and propose a granular evaluation of fuzzy models and present an augmentation of fuzzy models by forming information granules around numeric values of the parameters and constructions of the models. The concepts and algorithms of granular fuzzy models are discussed in the setting of Takagi-Sugeno rule-based architectures. We show how different protocols of forming and allocating information granules lead to the improvement of the granular performance of the models. Different from the standard numeric performance measure of fuzzy models coming in the form of the root mean squared error index, two performance measures are introduced that are pertinent to granular constructs, namely coverage and specificity. Furthermore, we propose a global indicator implied by these two measures, called an area under the curve, being computed for the characteristics of the granular model expressed in the coverage-specificity coordinates. A series of experimental studies is reported, which offers a comprehensive overview of the introduced performance measure criteria as well as the underlying realization of the granular fuzzy models. Xingchen Hu 0001, Witold Pedrycz, Xianmin Wang |
IEEE Trans. Fuzzy Syst. | 1 |
| 2015 | Comparative analysis of logic operators: A perspective of statistical testing and granular computing
Xingchen Hu 0001, Witold Pedrycz, Xianmin Wang |
Int. J. Approx. Reason. | 1 |