Xingchen Hu 0001

dblp:170/3965 · DBLP profile ↗
← Back
8ranked-venue papers in the field
1as first author
8since 2021 · last 2026
0000-0001-6879-5266ORCID · verified

Domains — venue-derived; a paper can count in several

Database Systems & Data Management · 4 (1 first)Knowledge Engineering, Semantic Web & Information Systems · 3Big Data, Cloud & Distributed Data Systems · 1
YearPublicationVenuePosition
2026 Contrastive and Dual Adversarial Representation Learning for Multi-View Clustering
abstract
Multi-View Clustering (MVC) has gained increasing attention due to its ability to effectively leverage the complementary information of multi-view data. Despite the success of existing MVC methods in many real-world applications, they often overlook the discrepancy of view-specific latent distribution and struggle to ensure the completeness of the multi-view data. To address these challenges and harness the powerful feature extraction capability of deep networks, we propose a novel Contrastive and Dual Adversarial Representation Learning method for Multi-view Clustering, termed as CDARL, to solve multi-view clustering problems with both complete and incomplete multi-view data. Specifically, CDARL employs alternating adversarial and contrastive learning to align the view-specific representations, driving them into the same semantic latent space to minimize the discrepancy in view-specific distributions. In addition, a consensus latent representation is learned by an adaptive fusion block that integrates information from multiple views. The consensus representation is further refined through adversarial learning modeling the transformation of the standard Gaussian distribution to the original data distribution. Moreover, the proposed method incorporates an imputation strategy designed to handle the incomplete multi-view data clustering task. This strategy utilizes both reconstructed samples and cross-view neighbors to impute missing views from the latent space and the original space, thereby preserving clustering information, which ensures the quality and feasibility of the imputed samples. Experimental results on six widely used datasets have verified the competitiveness of the proposed CDARL method against state-of-the-art methods in MVC problems with complete and incomplete multi-view data. Code is available athttps://github.com/xywy220/CDARL-MVC.
Yanwanyu Xi, Chang Tang, Junjie Huang 0001, Xingchen Hu 0001, Yuanyuan Liu 0004, Xinwang Liu 0002
IEEE Trans. Knowl. Data Eng.4
2025 Coherence mode: Characterizing local graph structural information for temporal knowledge graph
Yuehang Si, Xingchen Hu 0001, Qing Cheng 0004, Xinwang Liu 0002, Jincai Huang 0001
Inf. Sci.2
2025 Dynamic Ensemble Framework for Imbalanced Data Classification
abstract
Dynamic ensemble has significantly greater potential space to improve the classification of imbalanced data compared to static ensemble. However, dynamic ensemble schemes are far less successful than static ensemble methods in the imbalanced learning field. Through an in-depth analysis on the behavior characteristics of dynamic ensemble, we find that there are some important problems that need to be addressed to release the full potential of dynamic ensemble, including but not limited to, correcting the component classifiers’ bias towards the majority classes, increasing the proportions of the positive classifiers (i.e., the component classifiers making correct prediction) for difficult samples, and providing the accurate competence estimations on the hard-to-classify samples w.r.t the classifier pool. Inspired by these, we propose a Dynamic Ensemble Framework for imbalanced data classification (imDEF). imDEF first uses the data generation method OREM$\mathrm{_{G}}$to generate multiple artificial synthetic datasets, which have diverse class distributions by rebalancing the original imbalanced data. Based on each of such synthetic datasets, imDEF then utilizes a Classification Error-aware Self-Paced Sampling Ensemble (SPSE$\mathrm{_{CE}}$) method to gradually focus more on difficult samples, to create a low-biased classifier pool and increase the proportions of the positive classifiers for the difficult samples. Finally, imDEF constructs a referee system to achieve the competence estimations by leveraging an Ensemble Margin-aware Self-Paced Sampling Ensemble (SPSE$\mathrm{_{EM}}$) method. SPSE$\mathrm{_{EM}}$incrementally strengthens the learning of the hard-to-classify samples, so that the competent levels of component classifiers could be estimated accurately. Extensive experiments demonstrate the effectiveness of imDEF. The source codes have been made publicly available on GitHub.
Tuanfei Zhu, Xingchen Hu 0001, Xinwang Liu 0002, En Zhu, Xinzhong Zhu
IEEE Trans. Knowl. Data Eng.2
2024 Aligning the Representation of Knowledge Graph and Large Language Model for Causal Question Answering
abstract
Causal Question Answering (CQA) is essential for knowledge discovery, focusing on the intricate dynamics between events and entities without predefined contexts. Despite advancements of CQA models through Knowledge Graphs (KGs) and Pre-Trained Language Models (PLMs), existing approaches are hindered by knowledge conflict, insufficient capacity, and limitations in information fusion. Large Language Models (LLMs) have significantly improved natural language understanding and reasoning but often suffer from causal hallucinations. To address these challenges, we introduce KLop, a framework that aligns representations of Causal Knowledge Graph (CKG) and Large Language Models for CQA. KLop pre-trains a graph embedding model for entity embedding and uses a frozen LLM for text embedding. The main components of KLop are the descriptor module and the aligner module. The descriptor leverages descriptive texts generated by LLMs to create training data for knowledge alignment, while the aligner utilizes self-attention to train query tokens for modality alignment. Experiments on public CQA datasets validate that KLop outperforms various advanced baselines in reasoning accuracy, as well as achieving causal knowledge integration and joint reasoning.
Zefan Zeng, Qing Cheng 0004, Xingchen Hu 0001, Zhong Liu 0002, Jingke Shen, Yahao Zhang
IEEE Big Data3
2024 Discriminative embedded multi-view fuzzy C-means clustering for feature-redundant and incomplete data
Yan Li 0003, Xingchen Hu 0001, Tuanfei Zhu, Jiyuan Liu 0003, Xinwang Liu 0002, Zhong Liu 0002
Inf. Sci.2
2024 Application of Gradient Boosting in the Design of Fuzzy Rule-Based Regression Models
abstract
This study is devoted to the design of gradient boosted fuzzy rule-based models for regression problems. Fuzzy rule-based models are built on the basis of information granules formed in the input and output spaces whose structure involves a family of conditional ‘if-then’ statements. The architecture of fuzzy rule-based models contributes to the realization of a sound tradeoff between modeling accuracy and interpretability and computing overhead. Gradient boosting paradigm has emerged as a powerful learning method realized through sequentially fitting additive base learners to current residuals in the steepest descent way. However, surprisingly, studies on the design and analysis of gradient boosted fuzzy rule-based models are still lacking. In this study, fuzzy rule-based model is regarded as a base learner. Different loss functions and their influence on the performance of the final models are explored. We also thoroughly investigate an impact of the initial quality of the rule-based model (implied by the number of rules) on the process of gradient boosting. The performance of the proposed approach is illustrated by a series of experimental studies concerning synthetic and publicly available datasets.
Xingchen Hu 0001, Xiubin Zhu, Xinwang Liu 0002, Witold Pedrycz
IEEE Trans. Knowl. Data Eng.2
2023 Multi-View Fuzzy Classification With Subspace Clustering and Information Granules
abstract
Multi-view learning becomes increasingly attractive and promising because multimodal or multi-view data are commonly encountered in real-world applications. In this study, we develop a novel multi-view Takagi–Sugeno–Kang (TSK) fuzzy system framework to handle classification problems for such data. We propose an anchor and graph subspace clustering strategy to discover and represent the actual latent data distribution for each view separately. In this way, the discriminate anchors (landmarks) are learned to capture the main structure of the multi-view data. This strategy also provides a computationally efficient clustering algorithm with respect to the number of instances. These resulting anchors are formed as the prototypes of information granules (IGs) for fuzzy modeling. Then we construct an information-granule-based multi-view TSK fuzzy classification model inherited from the natural interpretability of fuzzy rule-based systems. Concretely, the relationship between the multi-view input and label output spaces is depicted by IGs-oriented fuzzy rules. The experimental studies involve various commonly used benchmark datasets, which indicate that our proposed method achieves comparable or better performance compared to the state-of-the-art algorithms.
Xingchen Hu 0001, Xinwang Liu 0002, Witold Pedrycz, Qing Liao 0001, Yinghua Shen, Yan Li 0003, Siwei Wang 0001
IEEE Trans. Knowl. Data Eng.1
2021 Fuzzy Rule-Based Models: A Design with Prototype Relocation and Granular Generalization
Yan Li 0003, Chao Chen 0017, Xingchen Hu 0001, Jindong Qin
Inf. Sci.3