VLDB 2026 Research / reviewers in the wild / expert
Yiqun Zhang 0006
dblp:125/5587-6
· DBLP profile ↗
57ranked-venue papers
14as first author
52since 2021 · last 2026
0000-0002-0328-987XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 32 · 11 first-author · 27 since 2021Graphics, computer vision, multimedia, augmented reality and games · 26 · 2 first-author · 25 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 7 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 5 since 2021Computer networks · 2 · 2 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Mask the Redundancy: Evolving Masking Representation Learning for Multivariate Time-Series ClusteringabstractMultivariate Time-Series (MTS) clustering discovers intrinsic grouping patterns of temporal data samples. Although time-series provide rich discriminative information, they also contain substantial redundancy, such as steady-state machine operation records and zero-output periods of solar power generation. Such redundancy diminishes the attention given to discriminative timestamps in representation learning, thus leading to performance bottlenecks in MTS clustering. Masking has been widely adopted to enhance the MTS representation, where temporal reconstruction tasks are designed to capture critical information from MTS. However, most existing masking strategies appear to be standalone preprocessing steps, isolated from the learning process, which hinders dynamic adaptation to the importance of clustering-critical timestamps. Accordingly, this paper proposes the Evolving-masked MTS Clustering (EMTC) method, whose model architecture comprises Importance-aware Variate-wise Masking (IVM) and Multi-Endogenous Views (MEV) generation modules. IVM adaptively guides the model in learning more discriminative representations for clustering, while the reconstruction and cluster-guided contrastive learning pathways enhance and connect the representation learning to clustering tasks. Extensive experiments on 15 benchmark datasets demonstrate the superiority of EMTC over eight SOTA methods, where the EMTC achieves an average improvement of 4.85% in F1-Score over the strongest baselines. Zexi Tan, Xiaopeng Luo, Yunlin Liu, Yiqun Zhang 0006 |
AAAI | 4 |
| 2026 | Break the Tie: Learning Cluster-Customized Category Relationships for Categorical Data ClusteringabstractCategorical attributes with qualitative values are ubiquitous in cluster analysis of real datasets. Unlike the Euclidean distance of numerical attributes, the categorical attributes lack well-defined relationships of their possible values (also called categories interchangeably), which hampers the exploration of compact categorical data clusters. Although most attempts are made for developing appropriate distance metrics, they typically assume a fixed topological relationship between categories when learning distance metrics, which limits their adaptability to varying cluster structures and often leads to suboptimal clustering performance. This paper, therefore, breaks the intrinsic relationship tie of attribute categories and learns customized distance metrics suitable for flexibly and accurately revealing various cluster distributions. As a result, the fitting ability of the clustering algorithm is significantly enhanced, benefiting from the learnable category relationships. Moreover, the learned category relationships are proved to be Euclidean distance metric-compatible, enabling a seamless extension to mixed datasets that include both numerical and categorical attributes. Comparative experiments on 12 real benchmark datasets with significance tests show the superior clustering accuracy of the proposed method with an average ranking of 1.25, which is significantly higher than the 5.21 ranking of the best-performing methods. Code and extended version with detailed proofs are provided online. Mingjie Zhao 0003, Zhanpei Huang, Yang Lu 0009, Mengke Li 0001, Yiqun Zhang 0006, Weifeng Su, Yiu-Ming Cheung |
AAAI | 5 |
| 2026 | HyReaL: Clustering Attributed Graph via Hyper-complex Space Representation Learning
Yang Lu 0009, Mengke Li 0001, Cuie Yang, Yiqun Zhang 0006, Yiu-Ming Cheung |
DASFAA (2) | 5 |
| 2026 | Bridging the Semantic Gap for Categorical Data Clustering via Large Language Models
Zihua Yang, Yiqun Zhang 0006, Yiu-Ming Cheung |
ICPR (2) | 3 |
| 2026 | One-Shot Federated Clustering of Non-Independent Completely Distributed DataabstractFederated Learning (FL) that extracts data knowledge while protecting the privacy of multiple clients has achieved remarkable results in distributed privacy-preserving IoT systems, including smart traffic flow monitoring, smart grid load balancing, and so on. Since most data collected from edge devices are unlabeled, unsupervised Federated Clustering (FC) is becoming increasingly popular for exploring pattern knowledge from complex distributed data. However, due to the lack of label guidance, the common Non-Independent and Identically Distributed (Non-IID) issue of clients have greatly challenged FC by posing the following problems: How to fuse pattern knowledge (i.e., cluster distribution) from Non-IID clients; How are the cluster distributions among clients related; and How does this relationship connect with the global knowledge fusion? In this paper, a more tricky but overlooked phenomenon in Non-IID is revealed, which bottlenecks the clustering performance of the existing FC approaches. That is, different clients could fragment a cluster, and accordingly, a more generalized Non-IID concept, i.e., Non-ICD (Non-Independent Completely Distributed), is derived. To tackle the above FC challenges, a new framework named GOLD (Global Oriented Local Distribution Learning) is proposed. GOLD first finely explores the potential incomplete local cluster distributions of clients, then uploads the distribution summarization to the server for global fusion, and finally performs local cluster enhancement under the guidance of the global distribution. Extensive experiments, including significance tests, ablation studies, scalability evaluations, qualitative results, etc., have been conducted to show the superiority of GOLD. Yiqun Zhang 0006, Shenghong Cai, Zihua Yang, Sen Feng, Yuzhu Ji, Haijun Zhang 0002 |
IEEE Internet Things J. | 1 |
| 2026 | Hierarchical Reference Sets for Robust Unsupervised Detection of Scattered and Clustered Outliersabstractreal-world IoT data analysis tasks, such as clustering and anomaly event detection, are unsupervised and highly susceptible to the presence of outliers. In addition to sporadic scattered outliers caused by factors such as faulty sensor readings, IoT systems often exhibit clustered outliers. These occur when multiple devices or nodes produce similar anomalous measurements, for instance, owing to localized interference, emerging security threats, or regional false alarms, forming micro-clusters. These clustered outliers can be easily mistaken for normal behavior because of their relatively high local density, thereby obscuring the detection of both scattered and contextual anomalies. To address this, we propose a novel outlier detection paradigm that leverages the natural neighboring relationships using graph structures. This facilitates multi-perspective anomaly evaluation by incorporating reference sets at both local and global scales derived from the graph. Our approach enables the effective recognition of scattered outliers without interference from clustered anomalies, whereas the graph structure simultaneously helps reflect and isolate clustered outlier groups. Extensive experiments, including comparative performance analysis, ablation studies, validation on downstream clustering tasks, and evaluation of hyperparameter sensitivity, demonstrate the efficacy of the proposed method. The source code is available at https://github.com/gordonlok/DROD. Yiqun Zhang 0006, Zexi Tan, Xiaopeng Luo, Yunlin Liu |
IEEE Internet Things J. | 1 |
| 2026 | Federated hierarchical clustering with automatic selection of optimal cluster numbers
Yue Zhang 0045, Chuanlong Qiu, Xinfa Liao, Yiqun Zhang 0006 |
Inf. Sci. | 4 |
| 2026 | Online Heterogeneous Feature SelectionabstractMany real-world datasets contain high-dimensional heterogeneous features, exhibiting complex and evolving distributions. The coexistence of high dimensionality and heterogeneity poses challenges for reliable feature selection and real-time analysis, while most existing feature selection solutions either assume that the features are of the same type or struggle to handle extremely high-dimensional features. Moreover, these methods are usually designed for static datasets, neglecting the dynamic capture of heterogeneous interfeature relationships in real-time environments. To address these challenges, we propose a new feature selection method called graph-unified adaptive decision boundary enhancement (GRADE) for online heterogeneous feature selection (OHFS). To provide a reliable foundation for evaluating feature subsets under dynamic and heterogeneous data streams, an incremental graph-unified metric (IGUM) is introduced. It mitigates information loss between heterogeneous features by leveraging graph structures to unify feature-value-level and interfeature-level relationships. With such a consistent relation measure, an adaptive density-guided neighborhood relation (ADNR) is proposed to assess the capability of selected feature subsets to classify samples. Since it dynamically captures prominent neighborhood regions, local decision boundaries can thus be precisely delineated. It turns out that GRADE can obtain a more concise feature subset while achieving competitive classification accuracy. Besides, GRADE is parameter-free and very efficient compared with state-of-the-art methods. Comprehensive experimental evaluations, including significance tests, ablation studies, efficiency evaluation, and case studies, have been conducted to verify the efficacy of GRADE. Yiqun Zhang 0006, Xinxi Chen, Lang Zhao, Yuzhu Ji, Peng Liu 0045, Yiu-Ming Cheung |
IEEE Trans. Cybern. | 1 |
| 2026 | NIDC: General Task Backbone for Neuroimaging Analysis via Interpretable Deep ClusteringabstractClustering techniques offer strong interpretability. However, they have significant limitations in the deep learning area due to their difficulty in capturing complex data structures, such as spatial and contextual information. This issue is especially pronounced in neuroimaging research, where high-dimensional and complex data greatly restricts the feature representation capability of clustering models. Hence, we propose a Neuroimaging Deep Clustering (NIDC) backbone network. We convert 3D neuroimagings into point sets and design clustering-based paradigms for context feature aggregation, feature interaction, and feature dispatching to enable deep feature extraction from the point sets. To better capture spatial information, we propose brain spatial relative position encoding, which assists the clustering paradigm in better understanding the anatomical structure of brain tissue and the positional relationships between different regions. Additionally, we design a sample center loss function to encourage tighter clustering of labels or voxels/feature points of the same class in the feature space, aiming to suppress both inter-class and intra-class similarities. Meanwhile, NIDC preserves the interpretability of traditional clustering techniques, allowing it to uncover relationships between brain regions and trace the decision-making process of each voxel. NIDC achieves highly competitive performance across multiple datasets in various downstream tasks, emerging as a new and practical solution for neuroimaging analysis. Code is available athttps://github.com/IMCTGD/NIDC. Jiayu Ye, An Zeng, Dan Pan 0001, Jingliang Zhao, Yiqun Zhang 0006, Yang Liu 0007 |
IEEE Trans. Multim. | 6 |
| 2025 | Asynchronous Federated Clustering with Unknown Number of ClustersabstractFederated Clustering (FC) is crucial to mining knowledge from unlabeled non-Independent Identically Distributed (non-IID) data provided by multiple clients while preserving their privacy. Most existing attempts learn cluster distributions at local clients, then securely pass the desensitized information to the server for aggregation. However, some tricky but common FC problems are still relatively unexplored, including the heterogeneity in terms of clients' communication capacity and the unknown number of proper clusters. To further bridge the gap between FC and real application scenarios, this paper first shows that the clients' communication asynchrony and unknown proper cluster numbers are complex coupling problems, and then proposes an Asynchronous Federated Cluster Learning (AFCL) method accordingly. It spreads the excessive number of seed points to clients as a learning medium and coordinates them across clients to form a consensus. To alleviate the distribution imbalance cumulated due to the unforeseen asynchronous uploading from the heterogeneous clients, we also design a balancing mechanism for seeds updating. As a result, the seeds gradually adapt to each other to reveal a proper number of clusters. Extensive experiments demonstrate the efficacy of AFCL. Yiqun Zhang 0006, Yang Lu 0009, Mengke Li 0001, Yiu-Ming Cheung |
AAAI | 2 |
| 2025 | DE3S: Dual-Enhanced Soft-Sparse Shape Learning for Medical Early Time Series ClassificationabstractEarly Time Series Classification (ETSC) is critical in time-sensitive medical applications such as sepsis, yet it presents an inherent trade-off between accuracy and earliness. This tradeoff arises from two core challenges: 1) models should effectively model inherently weak and noisy early-stage snippets, and 2) they should resolve the complex, dual requirement of simultaneously capturing local, subject-specific variations and overarching global temporal patterns. Existing methods struggle to overcome these underlying challenges, often forcing a severe compromise: sacrificing accuracy to achieve earliness, or vice-versa. We propose DE3S, a Dual-Enhanced Soft-Sparse Sequence Learning framework, which systematically solves these challenges. A dual enhancement mechanism is proposed to enhance the modeling of weak, early signals. Then, an attention-based patch module is introduced to preserve discriminative information while reducing noise and complexity. A dual-path fusion architecture is designed, using a sparse mixture of experts to model local, subject-specific variations. A multi-scale inception module is also employed to capture global dependencies. Experiments on six real-world medical datasets show the competitive performance of DE3S, particularly in early prediction windows. Ablation studies confirm the effectiveness of each component in addressing its targeted challenge. The source code is available here. Tao Xie 0011, Zexi Tan, Haoyi Xiao, Binbin Sun, Yiqun Zhang 0006 |
BIBM | 5 |
| 2025 | Mind the Gap: Confidence Discrepancy Can Guide Federated Semi-Supervised Learning Across Pseudo-MismatchabstractFederated Semi-Supervised Learning (FSSL) aims to leverage unlabeled data across clients with limited labeled data to train a global model with strong generalization ability. Most FSSL methods rely on consistency regularization with pseudo-labels, converting predictions from local or global models into hard pseudo-labels as supervisory signals. However, we discover that the quality of pseudo-label is largely deteriorated by data heterogeneity, an intrinsic facet of federated learning. In this paper, we study the problem of FSSL in-depth and show that (1) heterogeneity exacerbates pseudo-label mismatches, further degrading model performance and convergence, and (2) local and global models’ predictive tendencies diverge as heterogeneity increases. Motivated by these findings, we propose a simple and effective method called Semi-supervised Aggregation for Globally-Enhanced Ensemble (SAGE), that can flexibly correct pseudo-labels based on confidence discrepancies. This strategy effectively mitigates performance degradation caused by incorrect pseudo-labels and enhances consensus between local and global models. Experimental results demonstrate that SAGE outperforms existing FSSL methods in both performance and convergence. Our code is available at https://github.com/Jay-Codeman/SAGE. Xinyi Shang, Yiqun Zhang 0006, Yang Lu 0009, Chen Gong 0002, Jing-Hao Xue, Hanzi Wang |
CVPR | 3 |
| 2025 | FairFed++: Closing the Fairness Gap in Federated Learning Through Self-Evolving Clustered OptimizationabstractPerformance fairness in federated learning (FL) aims to ensure that the server treats all clients equitably, thereby encouraging the participation of low-performance clients and enhancing the generalization capabilities of the global model. However, due to client heterogeneity, conflicts may arise among the gradients of different clients during FL, suppressing performance fairness. Although current FL algorithms can achieve convergence, such conflicts lead to performance unfairness when reaching the optimal solution, thereby significantly restricting the generalization ability of the global model learned through FL. To address this issue, strategies such as reweighting and data augmentation have been proposed. However, these approaches often result in performance degradation for certain clients while striving for fairness. Recent studies have highlighted the potential of cluster federated learning (CFL) in achieving performance fairness. Nevertheless, the heavy reliance on a pre-specified number of clusters not only limits its adaptability but also increases the complexity of FL. Inspired by the principle of species evolution, where cells divide under specific internal conditions, we propose a novel FL method, namely FairFed++. Specifically, FairFed++ performs self-evolving clustered optimization, explicitly releasing the reliance on prior knowledge of clustering. By utilizing the accuracy variance within clusters as the splitting criterion, FairFed++ automatically determines the optimal clustering strategy in each round of FL communication until convergence. This approach dynamically adjusts the number of clusters during training without requiring manual intervention, thus improving both adaptability and practicability. Experiments conducted on six datasets demonstrate that FairFed++ achieves superior performance fairness while preserving the generalization ability of the global model. Zhixiang Fang, Baoyao Yang, Weide Zhan, Yanchao Tang, Yiqun Zhang 0006 |
ECAI | 5 |
| 2025 | Robust Qualitative Data Clustering via Learnable Multi-Metric Space FusionabstractUnderstanding categorical data with vague qualitative values by forming clusters is crucial in many data-driven AI fields. Compared with numerical data with its quantitative values embedded in well-defined Euclidean distance space, distances of the qualitative values are naturally unknown and are specially defined for certain data types or tasks. This paper, therefore, proposes a distance metric space fusion framework, which learns to fuse multiple distance metrics to form a statistical information-complete and prior knowledge-comprehensive metric for robust and accurate cluster analysis of qualitative data. To better serve various clustering tasks, the metric fusion objective is incorporated into the clustering objective through iterative learning. It turns out that the proposed method stably demonstrates superiority on various challenging real benchmark datasets. Extensive experiments including significance tests, ablation studies, etc. validate its efficacy. Source code of the proposed method is available at https://github.com/Sen-Feng/ICASSP-MSF/tree/main/CODE. Sen Feng, Mingjie Zhao 0003, Zhanpei Huang, Yuzhu Ji, Yiqun Zhang 0006, Yiu-Ming Cheung |
ICASSP | 5 |
| 2025 | SmartNet: One-shot Talking Head Synthesis via Subtle Motion and Appearance CompensationabstractOne-shot talking head synthesis aims to animate a source person’s portrait with driving video sequences. Recent facial keypoint-based methods have achieved remarkable animation performance and produced high-quality results. However, it remains challenging to perform cross-identity face reenactment by transferring subtle facial motions with correct geometry and appearance. To break the above limitations, in this paper, we propose a subtle motion compensation network to recover correct facial expressions by leveraging the decoupled 3D Morphable Model (3DMM) coefficient. In addition, to generate faithful animation results, a facial appearance feature memory bank is designed to learn accurate facial features and better recover the appearance. Experimental results have demonstrated that our proposed model can outperform state-of-the-art methods by generating faithful videos with correct subtle motion transfer and consistent identity preserving. Yuzhu Ji, An Zeng, Dan Pan 0001, Yiqun Zhang 0006, Haijun Zhang 0002 |
ICASSP | 5 |
| 2025 | Weighted Density for The Win: Accurate Subspace Density Clusteringabstractk-clustering typically struggles with the detection of irregular-distributed clusters due to the natural bias, while density clustering usually cannot well-adapt to different datasets and clustering tasks as it is not an oriented optimization process. This paper, therefore, proposes to perform density clustering in dynamically learned subspaces. To exploit the irregular-distributed clusters obtained by density clustering for the subspace determination, we design a new strategy to appropriately evaluate the importance of attributes. It turns out that the proposed Weighted Density-based Subspace Clustering (WDSC) algorithm inherits the unbiased merits of density clustering, and also upgrades the unlearning density clustering to be learnable under the subspace learning paradigm of k-clustering. A comprehensive evaluation including significance tests, ablation studies, qualitative comparisons, etc., shows the superiority of WDSC. Maixuan Peng, Yuyang Wu, Yang Lu 0009, Mengke Li 0001, Yiqun Zhang 0006, Yiu-Ming Cheung |
ICASSP | 5 |
| 2025 | PRO-VPT: Distribution-Adaptive Visual Prompt Tuning via Prompt RelocationabstractVisual prompt tuning (VPT), i.e., fine-tuning some lightweight prompt tokens, provides an efficient and effective approach for adapting pre-trained models to various downstream tasks. However, most prior art indiscriminately uses a fixed prompt distribution across different tasks, neglecting the importance of each block varying depending on the task. In this paper, we introduce adaptive distribution optimization (ADO) by tackling two key questions: (1) How to appropriately and formally define ADO, and (2) How to design an adaptive distribution strategy guided by this definition? Through empirical analysis, we first confirm that properly adjusting the distribution significantly improves VPT performance, and further uncover a key insight that a nested relationship exists between ADO and VPT. Based on these findings, we propose a new VPT framework, termed PRO-VPT (iterative Prompt RelOcation-based VPT), which adaptively adjusts the distribution built upon a nested optimization formulation. Specifically, we develop a prompt relocation strategy derived from this formulation, comprising two steps: pruning idle prompts from prompt-saturated blocks, followed by allocating these prompts to the most prompt-needed blocks. By iteratively performing prompt relocation and VPT, our proposal can adaptively learn the optimal prompt distribution in a nested optimization-based manner, thereby unlocking the full potential of VPT. Extensive experiments demonstrate that our proposal significantly outperforms advanced VPT methods, e.g., PRO-VPT surpasses VPT by 1.6 pp and 2.0 pp average accuracy, leading prompt-based methods to state-of-the-art performance on VTAB-1k and FGVC benchmarks. The code is available at https://github.com/ckshang/PRO-VPT. Chikai Shang, Mengke Li 0001, Yiqun Zhang 0006, Zhen Chen 0018, Jinlin Wu, Fangqing Gu, Yang Lu 0009, Yiu-Ming Cheung |
ICCV | 3 |
| 2025 | DeSAD: Density Clustering-Guided Streaming Data Anomaly Detection
Yue Zhang 0045, Xuchuang Ding, Gengwen Huang, Yiqun Zhang 0006 |
ICIC (9) | 6 |
| 2025 | SeqPose: An End-to-End Framework to Unify Single-frame and Video-based RGB Category-Level Pose EstimationabstractCategory-level object pose estimation is a longstanding and fundamental task crucial for augmented reality and robotic manipulation applications. Existing RGB-based approaches struggle with multi-stage settings and heavily rely on off-the-shelf techniques, such as object detectors, depth estimators, non-differentiable NOCS shape alignment, etc. Extra dependencies lead to the accumulation of errors and complicate the whole pipeline, limiting the deployment of these approaches in practical applications. This paper streamlined an end-to-end framework unifying the single-frame and video-based category-level pose estimation. Specifically, instead of explicitly introducing extra dependencies, the DINOv2 encoder and depth decoder, as robust semantic and geometric prior extractors, are leveraged to produce intra-frame hierarchical semantic and geometric features. A spatial-temporal sparse query network is developed to model the implicit correspondence and inter-frame correlations between a set of implicit 3D query anchors and intra-frame features. Finally, a pose prediction head is employed using the bipartite matching algorithm. Experimental results demonstrate that our model achieves state-of-the-art performance compared with RGB-based categorical pose estimation methods on the REAL275 and CAMERA25 datasets. Our code is available at https://andrewchiyz.github.io/vision.3dv.seqpose/. Yuzhu Ji, Mingshan Sun, Jianyang Shi, Xiaoke Jiang, Yiqun Zhang 0006, Haijun Zhang 0002 |
IJCAI | 5 |
| 2025 | EEG-TFNet: Spatiotemporal and Spectral Feature Integration for EEG-Based AD Detection
An Zeng, Zhao Guo, Dan Pan 0001, Yiqun Zhang 0006, Huisi Hong |
ISBRA (1) | 4 |
| 2025 | FATE: A Prompt-Tuning-Based Semi-Supervised Learning Framework for Extremely Limited Labeled DataabstractSemi-supervised learning (SSL) has achieved significant progress by leveraging both labeled data and unlabeled data. Existing SSL methods overlook a common real-world scenario when labeled data is extremely scarce, potentially as limited as a single labeled sample in the dataset. General SSL approaches struggle to train effectively from scratch under such constraints, while methods utilizing pre-trained models often fail to find an optimal balance between leveraging limited labeled data and abundant unlabeled data. To address this challenge, we propose Firstly Adapt, Then catEgorize (FATE), a novel SSL framework tailored for scenarios with extremely limited labeled data. At its core, the two-stage prompt tuning paradigm FATE exploits unlabeled data to compensate for scarce supervision signals, then transfers to downstream tasks. Concretely, FATE first adapts a pre-trained model to the feature distribution of downstream data using volumes of unlabeled samples in an unsupervised manner. It then applies an SSL method specifically designed for pre-trained models to complete the final classification task. FATE is designed to be compatible with both vision and vision-language pre-trained models. Extensive experiments demonstrate that FATE effectively mitigates challenges arising from the scarcity of labeled samples in SSL, achieving an average performance improvement of 33.74% across seven benchmarks compared to state-of-the-art SSL methods. Code is available at https://github.com/ganchi-huanggua/FATE.git. Hezhao Liu, Yang Lu 0009, Mengke Li 0001, Yiqun Zhang 0006, Shreyank N. Gowda, Chen Gong 0002, Hanzi Wang |
ACM Multimedia | 4 |
| 2025 | FedCD: A Hybrid Federated Learning Framework for Adaptive Training Under Data Heterogeneity
Weide Zhan, Baoyao Yang, Zhixiang Fang, Dongzhe Li, Yali Ma, Yiqun Zhang 0006 |
PRCV (9) | 6 |
| 2025 | MEET-Sepsis: Multi-Endogenous-View Enhanced Time-Series Representation Learning for Early Sepsis Prediction
Zexi Tan, Tao Xie 0011, Binbin Sun, Yiqun Zhang 0006, Yiu-Ming Cheung |
PRICAI (5) | 5 |
| 2025 | Towards Clustering of Incomplete Mixed-Attribute DataabstractABSTRACT Clustering analysis is one of the most important data mining and knowledge discovery tools in real applications. Since the widespread presence of missing values hampers clustering performance, missing values imputation becomes necessary for data pre‐processing. However, for the common datasets composed of both numerical and categorical attributes (also known as mixed‐attribute datasets), most existing imputation methods suffer from the following three limitations: (1) Only feasible for a certain type of attribute; (2) Encounter difficulties in considering the interdependence between different types of attributes; (3) Short in exploiting the information provided by the incomplete mix‐valued objects. As a result, the original data distribution can be ill‐restored, misleading the downstream clustering tasks. This paper therefore proposes a clustering‐imputation co‐learning method for incomplete mixed‐attribute datasets to address these issues. This method integrates imputation and clustering into one learning process, emphasising the interrelationships between mixed attributes during the imputation process and exploiting the information of incomplete objectsduring clustering. It turns out that appropriate recovery of the dataset and accurate clustering can be better achieved through a cross‐coupling manner. Experiments on various datasets validate the promising efficacy of the proposed method. Chuyao Zhang, Xinxi Chen, Zexi Tan, Fangqing Gu, Yuzhu Ji, Yiqun Zhang 0006 |
Expert Syst. J. Knowl. Eng. | 6 |
| 2025 | Learning unified distance metric for heterogeneous attribute data clustering
Yiqun Zhang 0006, Mingjie Zhao 0003, Yang Lu 0009, Yiu-Ming Cheung |
Expert Syst. Appl. | 1 |
| 2025 | SDENK: Unbiased subspace density-k-clusteringabstractClustering is one of the most important data analysis techniques, as it extracts knowledge without requiring data labels, making it crucial in many unsupervised application scenarios. However, conventional k -clustering struggles to detect irregularly distributed clusters owing to its inherent preference for convex clusters, while density-based clustering often lacks the ability to customise an appropriate metric space for different clustering tasks owing to the lack of a task-oriented optimisation process. To simultaneously address these biases in cluster shape and metric space, this paper proposes an unbiased hybrid framework to perform density-based clustering in subspaces. These subspaces are constructed through a newly developed attribute-weighted k -clustering paradigm. To exploit the irregularly distributed clusters obtained via density-based clustering for subspace learning, a novel strategy is designed to subdivide the clusters into compact sub-clusters, which are more suitable for evaluating attribute importance through k -clustering. As a result, the proposed subspace density K -clustering algorithm inherits the shape flexibility of density-based clustering and the metric adaptiveness of k -clustering. Moreover, the learnable design enables mutual optimisation between density clusters and subspaces, yielding robust and superior clustering performance across various datasets. Comprehensive evaluations, including comparative clustering performance evaluation, ablation studies, significance tests, noise-robustness evaluation, and hyper-parameter sensitivity studies, are conducted. Among 10 compared methods, the proposed Subspace DENsity-K-clustering (SDENK) achieves an average rank of 1.75 on 12 datasets in terms of the clustering accuracy metric ARI. • The merits of density and k -clustering are integrated, serving as a hybrid general framework for unbiased clustering. • A novel re-clustering mechanism is proposed to subdivide the density clusters for more appropriately determining the subspace. • The proposed unbiased clustering algorithm is suitable for datasets with various cluster distributions. Mingjie Zhao 0003, Zexi Tan, Yiqun Zhang 0006, Yiu-Ming Cheung |
Neurocomputing | 5 |
| 2025 | Categorical Data Clustering via Value Order Estimated Distance Metric LearningabstractClustering is a popular machine learning technique for data mining that can process and analyze datasets to automatically reveal sample distribution patterns. Since the ubiquitous categorical data naturally lack a well-defined metric space such as the Euclidean distance space of numerical data, the distribution of categorical data is usually under-represented, and thus valuable information can be easily twisted in clustering. This paper, therefore, introduces a novel order distance metric learning approach to intuitively represent categorical attribute values by learning their optimal order relationship and quantifying their distance in a line similar to that of the numerical attributes. Since subjectively created qualitative categorical values involve ambiguity and fuzziness, the order distance metric is learned in the context of clustering. Accordingly, a new joint learning paradigm is developed to alternatively perform clustering and order distance metric learning with low time complexity and a guarantee of convergence. Due to the clustering-friendly order learning mechanism and the homogeneous ordinal nature of the order distance and Euclidean distance, the proposed method achieves superior clustering accuracy on categorical and mixed datasets. More importantly, the learned order distance metric greatly reduces the difficulty of understanding and managing the non-intuitive categorical data. Experiments with ablation studies, significance tests, case studies, etc., have validated the efficacy of the proposed method. The source code is available at https://github.com/csmjzhao/OCL_Source_Code. Yiqun Zhang 0006, Mingjie Zhao 0003, Hong Jia, Mengke Li 0001, Yang Lu 0009, Yiu-Ming Cheung |
Proc. ACM Manag. Data | 1 |
| 2025 | CmdVIT: A Voluntary Facial Expression Recognition Model for Complex Mental DisordersabstractFacial Expression Recognition (FER) is a critical method for evaluating the emotional states of patients with mental disorders, playing a significant role in treatment monitoring. However, due to privacy constraints, facial expression data from patients with mental disorders is severely limited. Additionally, the more complex inter-class and intra-class similarities compared to healthy individuals make accurate recognition of facial expressions challenging. Therefore, we propose a Voluntary Facial Expression Mimicry (VFEM) experiment, which collected facial expression data from schizophrenia, depression, and anxiety. This experiment establishes the first dataset designed for facial expression recognition tasks exclusively composed of patients with mental disorders. Simultaneously, based on VFEM, we propose a Vision Transformer FER model tailored for Complex mental disorder patients (CmdVIT). CmdVIT integrates crucial facial expression features through both explicit and implicit mechanisms, including explicit visual center positional encoding and implicit sparse attention center loss function. These two key components enhance positional information and minimize the facial feature space distance between conventional attention and critical attention, effectively suppressing inter-class and intra-class similarities. In various FER tasks for different mental disorders in VFEM, CmdVIT achieves more competitive performance compared to contemporary benchmark models. Our works are available at https://github.com/yjy-97/CmdVIT. Jiayu Ye, Yanhong Yu, Qingxiang Wang, Guolong Liu, An Zeng, Yiqun Zhang 0006, Yang Liu 0007, Yunshao Zheng |
IEEE Trans. Image Process. | 7 |
| 2025 | Allosteric Feature Collaboration for Model-Heterogeneous Federated LearningabstractAlthough federated learning (FL) has achieved outstanding results in privacy-preserved distributed learning, the setting of model homogeneity among clients restricts its wide application in practice. This article investigates a more general case, namely, model-heterogeneous FL (M-hete FL), where client models are independently designed and can be structurally heterogeneous. M-hete FL faces new challenges in collaborative learning because the parameters of heterogeneous models could not be directly aggregated. In this article, we propose a novel allosteric feature collaboration (AlFeCo) method, which interchanges knowledge across clients and collaboratively updates heterogeneous models on the server. Specifically, an allosteric feature generator is developed to reveal task-relevant information from multiple client models. The revealed information is stored in the client-shared and client-specific codes. We exchange client-specific codes across clients to facilitate knowledge interchange and generate allosteric features that are dimensionally variable for model updates. To promote information communication between different clients, a dual-path (model-model and model-prediction) communication mechanism is designed to supervise the collaborative model updates using the allosteric features. Client models are fully communicated through the knowledge interchange between models and between models and predictions. We further provide theoretical evidence and convergence analysis to support the effectiveness of AlFeCo in M-hete FL. The experimental results show that the proposed AlFeCo method not only performs well on classical FL benchmarks but also is effective in model-heterogeneous federated antispoofing. Our codes are publicly available at https://github.com/ybaoyao/AlFeCo. Baoyao Yang, Pong C. Yuen, Yiqun Zhang 0006, An Zeng |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2025 | Learning Self-Growth Maps for Fast and Accurate Imbalanced Streaming Data ClusteringabstractStreaming data clustering is a popular research topic in data mining and machine learning. Since streaming data is usually analyzed in data chunks, it is more susceptible to encountering the dynamic cluster imbalance issue. That is, the imbalance ratio (IR) of clusters changes over time, which can easily lead to fluctuations in either the accuracy or the efficiency of streaming data clustering. Therefore, an accurate and efficient streaming data clustering approach is proposed to adapt to the drifting and imbalanced cluster distributions. We first design a self-growth map (SGM) that can automatically arrange neurons on demand according to local distribution, and thus achieve fast and incremental adaptation to the streaming distributions. Since SGM allocates an excess number of density-sensitive neurons to describe the global distribution, it can avoid missing small clusters among imbalanced distributions. We also propose a fast hierarchical merging (HM) strategy to combine the neurons that break up the relatively large clusters. It exploits the maintained SGM to quickly retrieve the intracluster distribution pairs for merging, which circumvents the most laborious global searching. It turns out that the proposed SGM can incrementally adapt to the distributions of new chunks, and the self-growth map-guided hierarchical merging for the imbalanced data clustering (SOHI) approach can quickly explore a true number of imbalanced clusters. Extensive experiments demonstrate that SOHI can efficiently and accurately explore cluster distributions for streaming data. Yiqun Zhang 0006, Sen Feng, Zexi Tan, Xiaopeng Luo, Yuzhu Ji, Yiu-Ming Cheung |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | MGOD: Multi-Granular Outlier Detection with Clustlier AnalysisabstractUnsupervised Outlier Detection (UOD) is crucial for the analysis of biomedical and health data with undesirable outliers. However, the complex distribution of real data often brings difficulties to UOD where the "masking effect", i.e., only a small number of densely distributed outliers (also called clustliers) can collectively mask themselves from being detected, is particularly challenging. Another difficulty derived from this is how to distinguish clustliers from small clusters. Therefore, we propose a novel Multi-Granular Outlier Detector (MGOD). It first partitions the dataset into subsets with natural neighbor topological relationships to circumvent the non-trivial neighbor range setting. Then it effectively detects both clustliers and isolated samples (also called scatliers) based on a newly designed anomaly score. The score comprehensively takes into account the density and connectivity of samples to reflect different extents and types of abnormality. It turns out that MGOD is accurate and highly interpretable. The performance of MGOD is also robust to the involved hyper-parameters, which are easy to set. Comprehensive evaluations have been conducted to compare seven counterparts on 15 datasets, most of which are biomedical datasets. The results of significance tests confirm the effectiveness and superiority of MGOD. The source code is opened at https://anonymous.4open.science/r/MGOD-C531. Qingsheng Chen, Mingjie Zhao 0003, Yuzhu Ji, Xiaopeng Luo, Yiqun Zhang 0006, Yue Zhang 0045 |
BIBM | 5 |
| 2024 | Efficient Topology-Driven Clustering for Imbalanced Streaming Biomedical Data AnalysisabstractClustering drifting data is common in the field of biomedical data analysis. Data chunks collected at different periods often exhibit clusters with significantly different sizes, and drifting distributions of clusters also appear frequently. We call such composite phenomenon imbalance-drifting, which can severely impact the accuracy and efficiency of cluster analysis. Therefore, we propose a topology-representation-based clustering paradigm, which first learns an informative global data representation in a self-organizing manner to obtain a map with nested representative data points. Then fast and accurate clustering is facilitated by quickly retrieving similar data points according to the topology. As the constructed Self-Organizing Map (SOM) is exploited for informative representation, micro partition, and quick merging, to achieve advanced clustering under imbalance-drifting, the proposed approach is thus called Tri-Squeezing SOM for Clustering (TSSC). It turns out that TSSC significantly reduces the time complexity for clustering an n-scale imbalance-streaming data without sacrificing accuracy. Moreover, TSSC can automatically determine the number of clusters k, and features interpretability and hyper-parameter robustness. Extensive results on both biomedical datasets and synthetic datasets verify the superiority of TSSC. Xiaopeng Luo, Yiqun Zhang 0006, Yuzhu Ji, Peng Liu 0045, Taoting Xiao |
BIBM | 2 |
| 2024 | CPNet: 3D Semantic Relation and Geometry Context Prior Network for Multi-Organ SegmentationabstractAutomatic multi-organ segmentation of the abdominal region is a critical yet challenging task in computer-aided medical image analysis. Recent advances in CNN- and Transformer-based encoder-decoder models tend to implicitly learn context features by using enhanced effective receptive fields to capture local and global range dependencies. However, due to the complex anatomical structure, those models cannot recover the anatomical topology properly and result in broken organs with inaccurate semantic labels. Therefore, in this paper, by considering the anatomy priors of multi-organs, we propose a Context Prior Network, namely CPNet, which integrates the 3D context semantic relations and geometry priors as explicit anatomical constraints. Specifically, a Semantic Relation Prior Propagation (SRPP) module is designed to propagate the semantic relations between voxels progressively. Moreover, a Multiple Context Prior Prediction (MCPP) module is adopted to preserve the accurate shape and topology by recovering 3D contours and surface normal. Experimental results demonstrate our proposed model outperforms state-of-the-art models for multi-organ segmentation on Abdomen CT and MRI datasets, especially for recovering organs with correct semantic labels and anatomical structures. Yuzhu Ji, Mingshan Sun, Yiqun Zhang 0006, Haijun Zhang 0002 |
ECAI | 3 |
| 2024 | Learning Order Forest for Qualitative-Attribute Data ClusteringabstractClustering is a fundamental approach to understanding data patterns, wherein the intuitive Euclidean distance space is commonly adopted. However, this is not the case for implicit cluster distributions reflected by qualitative attribute values, e.g., the nominal values of attributes like symptoms, marital status, etc. This paper, therefore, discovered a tree-like distance structure to flexibly represent the local order relationship among intra-attribute qualitative values. That is, treating a value as the vertex of the tree allows to capture rich order relationships among the vertex value and the others. To obtain the trees in a clustering-friendly form, a joint learning mechanism is proposed to iteratively obtain more appropriate tree structures and clusters. It turns out that the latent distance space of the whole dataset can be well-represented by a forest consisting of the learned trees. Extensive experiments demonstrate that the joint learning adapts the forest to the clustering task to yield accurate results. Comparisons of 10 counterparts on 12 real benchmark datasets with significance tests verify the superiority of the proposed method. Source code of the proposed method is available at [39]. Mingjie Zhao 0003, Sen Feng, Yiqun Zhang 0006, Mengke Li 0001, Yang Lu 0009, Yiu-Ming Cheung |
ECAI | 3 |
| 2024 | Robust Categorical Data Clustering Guided by Multi-Granular Competitive LearningabstractData set composed of categorical features is very common in big data analysis tasks. Since categorical features are usually with a limited number of qualitative possible values, the nested granular cluster effect is prevalent in the implicit discrete distance space of categorical data. That is, data objects frequently overlap in space or subspace to form small compact clusters, and similar small clusters often form larger clusters. However, the distance space cannot be well-defined like the Euclidean distance due to the qualitative categorical data values, which brings great challenges to the cluster analysis of categorical data. In view of this, we design a Multi-Granular Competitive Penalization Learning (MGCPL) algorithm to allow potential clusters to interactively tune themselves and converge in stages with different numbers of naturally compact clusters. To leverage MGCPL, we also propose a Cluster Aggregation strategy based on MGCPL Encoding (CAME) to first encode the data objects according to the learned multi-granular distributions, and then perform final clustering on the embeddings. It turns out that the proposed MGCPL-guided Categorical Data Clustering (MCDC) approach is competent in automatically exploring the nested distribution of multi-granular clusters and highly robust to categorical data sets from various domains. Benefiting from its linear time complexity, MCDC is scalable to large-scale data sets and promising in pre-partitioning data sets or compute nodes for boosting distributed computing. Extensive experiments with statistical evidence demonstrate its superiority compared to state-of-the-art counterparts on various real public data sets. Shenghong Cai, Yiqun Zhang 0006, Xiaopeng Luo, Yiu-Ming Cheung, Hong Jia, Peng Liu 0045 |
ICDCS | 2 |
| 2024 | Towards Unbiased Minimal Cluster Analysis of Categorical-and-Numerical Attribute Data
Xiaopeng Luo, Qingsheng Chen, Yiqun Zhang 0006, Yiu-Ming Cheung |
ICPR (2) | 5 |
| 2024 | Clustering by Learning the Ordinal Relationships of Qualitative Attribute ValuesabstractIn many real-world clustering tasks, data objects are described by both quantitative and qualitative attributes. Attributes with semantically ordered qualitative values are very common and are usually coded according to their order (i.e., consecutive integers) for clustering. However, semantic order is not always globally interdependent with a certain clustering task. An intuitive case is that level of income (attribute) is not always positively correlated with the level of mental health (label). Using mismatched order surely forms a bottleneck to clustering performance, and conversely, the unsupervised clustering process prevents understanding of "true" order. Therefore, we proposed a novel learning paradigm to tune the value order. More specifically, we adjust the intra-attribute orders, and let this process learn mutually with object clustering, thus bridging the gap between value order and clustering task. To the best of our knowledge, this is the first attempt to learn ordinal relationships among qualitative attribute values. Extensive experiments with significance tests show that our method outperforms the existing relevant clustering approaches on qualitative attribute data. Yiqun Zhang 0006, Yang Lu 0009, Mengke Li 0001, Yiu-Ming Cheung |
IJCNN | 3 |
| 2024 | QGRL: Quaternion Graph Representation Learning for Heterogeneous Feature Data ClusteringabstractClustering is one of the most commonly used techniques for unsupervised data analysis.As real data sets are usually composed of numerical and categorical features that are heterogeneous in nature, the heterogeneity in the distance metric and feature coupling prevents deep representation learning from achieving satisfactory clustering accuracy.Currently, supervised Quaternion Representation Learning (QRL) has achieved remarkable success in efficiently learning informative representations of coupled features from multiple views derived endogenously from the original data.To inherit the advantages of QRL for unsupervised heterogeneous feature representation learning, we propose a deep QRL model that works in an encoder-decoder manner.To ensure that the implicit couplings of heterogeneous feature data can be well characterized by representation learning, a hierarchical coupling encoding strategy is designed to convert the data set into an attributed graph to be the input of QRL.We also integrate the clustering objective into the model training to facilitate a joint optimization of the representation and clustering.Extensive experimental evaluations illustrate the superiority of the proposed Quaternion Graph Representation Learning (QGRL) method in terms of clustering accuracy and robustness to various data sets composed of arbitrary combinations of numerical and categorical features.The source code is opened at https://github.com/Juny-Chen/QGRL.git. Yuzhu Ji, Yiqun Zhang 0006, Yiu-Ming Cheung |
KDD | 4 |
| 2024 | Improving Visual Prompt Tuning by Gaussian Neighborhood Minimization for Long-Tailed Visual RecognitionabstractLong-tailed visual recognition has received increasing attention recently. Despite fine-tuning techniques represented by visual prompt tuning (VPT) achieving substantial performance improvement by leveraging pre-trained knowledge, models still exhibit unsatisfactory generalization performance on tail classes. To address this issue, we propose a novel optimization strategy called Gaussian neighborhood minimization prompt tuning (GNM-PT), for VPT to address the long-tail learning problem. We introduce a novel Gaussian neighborhood loss, which provides a tight upper bound on the loss function of data distribution, facilitating a flattened loss landscape correlated to improved model generalization. Specifically, GNM-PT seeks the gradient descent direction within a random parameter neighborhood, independent of input samples, during each gradient update. Ultimately, GNM-PT enhances generalization across all classes while simultaneously reducing computational overhead. The proposed GNM-PT achieves state-of-the-art classification accuracies of 90.3%, 76.5%, and 50.1% on benchmark datasets CIFAR100-LT (IR 100), iNaturalist 2018, and Places-LT, respectively. The source code is available at https://github.com/Keke921/GNM-PT. Mengke Li 0001, Yang Lu 0009, Yiqun Zhang 0006, Yiu-Ming Cheung, Hui Huang 0004 |
NeurIPS | 4 |
| 2024 | FedGC: Federated Learning on Non-IID Data via Learning from Good Clients
Ting Cui, Yiqun Zhang 0006 |
PRCV (1) | 4 |
| 2024 | FedHC: Learning Imbalanced Clusters via Federated Hierarchical Clustering
Yue Zhang 0045, Xinfa Liao, Qingsheng Chen, Haotian Wu 0009, Yiqun Zhang 0006 |
PRCV (1) | 5 |
| 2024 | MAD-Former: A Traceable Interpretability Model for Alzheimer's Disease Recognition Based on Multi-Patch AttentionabstractThe integration of structural magnetic resonance imaging (sMRI) and deep learning techniques is one of the important research directions for the automatic diagnosis of Alzheimer's disease (AD). Despite the satisfactory performance achieved by existing voxel-based models based on convolutional neural networks (CNNs), such models only handle AD-related brain atrophy at a single spatial scale and lack spatial localization of abnormal brain regions based on model interpretability. To address the above limitations, we propose a traceable interpretability model for AD recognition based on multi-patch attention (MAD-Former). MAD-Former consists of two parts: recognition and interpretability. In the recognition part, we design a 3D brain feature extraction network to extract local features, followed by constructing a dual-branch attention structure with different patch sizes to achieve global feature extraction, forming a multi-scale spatial feature extraction framework. Meanwhile, we propose an important attention similarity position loss function to assist in model decision-making. The interpretability part proposes a traceable method that can obtain a 3D ROI space through attention-based selection and receptive field tracing. This space encompasses key brain tissues that influence model decisions. Experimental results reveal the significant role of brain tissues such as the Fusiform Gyrus (FuG) in AD recognition. MAD-Former achieves outstanding performance in different tasks on ADNI and OASIS datasets, demonstrating reliable model interpretability. Jiayu Ye, An Zeng, Dan Pan 0001, Yiqun Zhang 0006, Jingliang Zhao, Qiuping Chen, Yang Liu 0007 |
IEEE J. Biomed. Health Informatics | 4 |
| 2023 | Selecting Heterogeneous Features Based on Unified Density-Guided Neighborhood Relation for Complex Biomedical Data AnalysisabstractBiomedical big data are usually high dimensional and collected in the form of a continuous influx of new features. Online Feature Selection (OFS) is a promising way to manage and analyze such data, as OFS circumvents the huge computation cost brought by simultaneously considering all the features, and can also dynamically maintain a distribution-fitting feature subset on the fly. However, almost all the OFS solutions are based on a naive premise that all features are of the same type, overlooking the fact that real biomedical data set usually consists of heterogeneous numerical and categorical features. This paper therefore proposes a new approach to Online Heterogeneous Feature Selection (OHFS), which dynamically maintains a feature subset that maximizes the number of neighborhood sets where all the objects within each neighborhood set are of the same class. To appropriately partition the objects into neighborhood sets, a density-guided relation is proposed, which adaptively forms non-overlapping neighborhood sets by detecting spatially compact objects. A unified density measure is also presented to avoid information loss in processing heterogeneous features. It turns out that the proposed approach features parameter- free, interpretability, and efficiency. It is capable of maintaining a concise feature subset while receiving any type of feature. Extensive experimental evaluations demonstrate its superiority. Lang Zhao, Yiqun Zhang 0006, Xiaopeng Luo, Yue Zhang 0045, Yiu-Ming Cheung, Kangshun Li |
BIBM | 2 |
| 2023 | Time-Series Data Imputation via Realistic Masking-Guided Tri-Attention Bi-GRUabstractTime series data with missing values are ubiquitous in real applications due to various unforeseen faults during data generation, storage, and transmission. Time-Series Data Imputation (TSDI) is thus crucial to many temporal data analysis tasks. However, existing works usually consider only one of the following two issues: (1) intra-feature temporal dependency, and (2) inter-feature correlation, leading to the overlook of complex coupling information in imputation. To achieve more accurate TDSI, we design a novel imputation model called TABiG, which delicately preserves the short-term, long-term, and inter-feature dependencies by attention mechanisms in a delay error-reduced bi-directional architecture. That is, it leverages GRU to model short-term temporal dependencies and adopts self-attention mechanisms hierarchically to capture long-term temporal dependencies and inter-feature correlations. The multiple self-attention mechanisms are nested in a bi-directional structure to alleviate the problem of delay errors in RNN-like structures. To facilitate model training with higher generalization, a masking strategy that mimics various extreme real missing situations beyond the simple random ones has been adopted for generating self-supervised learning tasks. Comprehensive experiments demonstrate that TABiG significantly outperforms most state-of-the-art imputation counterparts. Complementary results and source code can be accessed at https://github.com/Zhang2112105189/TABiG Yiqun Zhang 0006, An Zeng, Dan Pan 0001, Yuzhu Ji |
ECAI | 2 |
| 2023 | CFNet: A Coarse-to-Fine Framework for Coronary Artery Segmentation
Shiting He, Yuzhu Ji, Yiqun Zhang 0006, An Zeng, Dan Pan 0001 |
PRCV (5) | 3 |
| 2023 | Learning Hierarchical Representations in Temporal and Frequency Domains for Time Series Forecasting
Yiqun Zhang 0006, An Zeng, Dan Pan 0001 |
PRCV (9) | 2 |
| 2023 | Unsupervised Concept Drift Detection via Imbalanced Cluster Discriminator Learning
Mingjie Zhao 0003, Yiqun Zhang 0006, Yuzhu Ji, Yang Lu 0009 |
PRCV (3) | 2 |
| 2023 | Graph-Based Dissimilarity Measurement for Cluster Analysis of Any-Type-Attributed DataabstractHeterogeneous attribute data composed of attributes with different types of values are quite common in a variety of real-world applications. As data annotation is usually expensive, clustering has provided a promising way for processing unlabeled data, where the adopted similarity measure plays a key role in determining the clustering accuracy. However, it is a very challenging task to appropriately define the similarity between data objects with heterogeneous attributes because the values from heterogeneous attributes are generally with very different characteristics. Specifically, numerical attributes are with quantitative values, while categorical attributes are with qualitative values. Furthermore, categorical attributes can be categorized into nominal and ordinal ones according to the order information of their values. To circumvent the awkward gap among the heterogeneous attributes, this article will propose a new dissimilarity metric for cluster analysis of such data. We first study the connections among the heterogeneous attributes and build graph representations for them. Then, a metric is proposed, which computes the dissimilarities between attribute values under the guidance of the graph structures. Finally, we develop a new k -means-type clustering algorithm associated with this proposed metric. It turns out that the proposed method is competent to perform cluster analysis of datasets composed of an arbitrary combination of numerical, nominal, and ordinal attributes. Experimental results show its efficacy in comparison with its counterparts. Yiqun Zhang 0006, Yiu-Ming Cheung |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2022 | Heterogeneous Drift Learning: Classification of Mix-Attribute Data with Concept DriftsabstractAs many real data sets (e.g., social, financial, and medical data sets) are successively generated in evolution with the ever-changing environment, classification for data stream with concept drift attracts increasing attention in the fields of machine learning and data mining. However, to the best of our knowledge, existing works mainly consider the concept drift issue while ignoring another common characteristic of real data, i.e., existence of awkward heterogeneity caused by mixture of numerical and categorical attributes. It is worth noting that tackling both the concept drift and heterogeneity problems together is exponentially more challenging than dealing with only one of them. This paper, therefore, proposes an ensemble learning approach for the classification of numerical-and-categorical-attribute data (also called mixed data hereinafter) under concept drift. We first design a unified metric to appropriately address the heterogeneity of numerical and categorical attributes. Then a base classifier that can appropriately fuse the information provided by the heterogeneous attributes is formed accordingly. Furthermore, to make the classification adapt to the complex concept drifts demonstrated on the heterogeneous attributes, two types of base classifier ensembles are dynamically learned on the fly. Experimental results on various real mixed data sets with concept drifts demonstrate the efficacy of the proposed method. Lang Zhao, Yiqun Zhang 0006, Yuzhu Ji, An Zeng, Fangqing Gu, Xiaopeng Luo |
DSAA | 2 |
| 2022 | Het2Hom: Representation of Heterogeneous Attributes into Homogeneous Concept Spaces for Categorical-and-Numerical-Attribute Data ClusteringabstractData sets composed of a mixture of categorical and numerical attributes (also called mixed data hereinafter) are common in real-world cluster analysis. However, insightful analysis of such data under an unsupervised scenario using clustering is extremely challenging because the information provided by the two different types of attributes is heterogeneous, being at different concept hierarchies. That is, the values of a categorical attribute represent a set of different concepts (e.g., professor, lawyer, and doctor of the attribute "occupation"), while the values of a numerical attribute describe the tendencies toward two different concepts (e.g., low and high of the attribute "income"). To appropriately use such heterogeneous information in clustering, this paper therefore proposes a novel attribute representation learning method called Het2Hom, which first converts the heterogeneous attributes into a homogeneous form, and then learns attribute representations and data partitions on such a homogeneous basis. Het2Hom features low time complexity and intuitive interpretability. Extensive experiments show that Het2Hom outperforms the state-of-the-art counterparts. Yiqun Zhang 0006, Yiu-Ming Cheung, An Zeng |
IJCAI | 1 |
| 2022 | Learnable Weighting of Intra-Attribute Distances for Categorical Data Clustering with Nominal and Ordinal AttributesabstractThe success of categorical data clustering generally much relies on the distance metric that measures the dissimilarity degree between two objects. However, most of the existing clustering methods treat the two categorical subtypes, i.e., nominal and ordinal attributes, in the same way when calculating the dissimilarity without considering the relative order information of the ordinal values. Moreover, there would exist interdependence among the nominal and ordinal attributes, which is worth exploring for indicating the dissimilarity. This paper will therefore study the intrinsic difference and connection of nominal and ordinal attribute values from a perspective akin to the graph. Accordingly, we propose a novel distance metric to measure the intra-attribute distances of nominal and ordinal attributes in a unified way, meanwhile preserving the order relationship among ordinal values. Subsequently, we propose a new clustering algorithm to make the learning of intra-attribute distance weights and partitions of data objects into a single learning paradigm rather than two separate steps, whereby circumventing a suboptimal solution. Experiments show the efficacy of the proposed algorithm in comparison with the existing counterparts. Yiqun Zhang 0006, Yiu-Ming Cheung |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | A New Distance Metric Exploiting Heterogeneous Interattribute Relationship for Ordinal-and-Nominal-Attribute Data ClusteringabstractOrdinal attribute has all the common characteristics of a nominal one but it differs from the nominal one by having naturally ordered possible values (also called categories interchangeably). In clustering analysis tasks, categorical data composed of both ordinal and nominal attributes (also called mixed-categorical data interchangeably) are common. Under this circumstance, existing distance and similarity measures suffer from at least one of the following two drawbacks: 1) directly treat ordinal attributes as nominal ones, and thus ignore the order information from them and 2) suppose all the attributes are independent of each other, measure the distance between two categories from a target attribute without considering the valuable information provided by the other attributes that correlate with the target one. These two drawbacks may twist the natural distances of attributes and further lead to unsatisfactory clustering results. This article, therefore, presents an entropy-based distance metric that quantifies the distance between categories by exploiting the information provided by different attributes that correlate with the target one. It also preserves the order relationship among ordinal categories during the distance measurement. Since attributes are usually correlated in different degrees, we also define the interdependence between different types of attributes to weight their contributions in forming distances. The proposed metric overcomes the two above-mentioned drawbacks for mixed-categorical data clustering. More important, it conceptually unifies the distances of ordinal and nominal attributes to avoid information loss during clustering. Moreover, it is parameter free, and will not bring extra computational cost compared to the existing state-of-the-art counterparts. Extensive experiments show the superiority of the proposed distance metric. Yiqun Zhang 0006, Yiu-Ming Cheung |
IEEE Trans. Cybern. | 1 |
| 2020 | An Ordinal Data Clustering Algorithm with Automated Distance LearningabstractClustering ordinal data is a common task in data mining and machine learning fields. As a major type of categorical data, ordinal data is composed of attributes with naturally ordered possible values (also called categories interchangeably in this paper). However, due to the lack of dedicated distance metric, ordinal categories are usually treated as nominal ones, or coded as consecutive integers and treated as numerical ones. Both these two common ways will roughly define the distances between ordinal categories because the former way ignores the order relationship and the latter way simply assigns identical distances to different pairs of adjacent categories that may have intrinsically unequal distances. As a result, they may produce unsatisfactory ordinal data clustering results. This paper, therefore, proposes a novel ordinal data clustering algorithm, which iteratively learns: 1) The partition of ordinal dataset, and 2) the inter-category distances. To the best of our knowledge, this is the first attempt to dynamically adjust inter-category distances during the clustering process to search for a better partition of ordinal data. The proposed algorithm features superior clustering accuracy, low time complexity, fast convergence, and is parameter-free. Extensive experiments show its efficacy. Yiqun Zhang 0006, Yiu-Ming Cheung |
AAAI | 1 |
| 2020 | A Unified Entropy-Based Distance Metric for Ordinal-and-Nominal-Attribute Data ClusteringabstractOrdinal data are common in many data mining and machine learning tasks. Compared to nominal data, the possible values (also called categories interchangeably) of an ordinal attribute are naturally ordered. Nevertheless, since the data values are not quantitative, the distance between two categories of an ordinal attribute is generally not well defined, which surely has a serious impact on the result of the quantitative analysis if an inappropriate distance metric is utilized. From the practical perspective, ordinal-and-nominal-attribute categorical data, i.e., categorical data associated with a mixture of nominal and ordinal attributes, is common, but the distance metric for such data has yet to be well explored in the literature. In this paper, within the framework of clustering analysis, we therefore first propose an entropy-based distance metric for ordinal attributes, which exploits the underlying order information among categories of an ordinal attribute for the distance measurement. Then, we generalize this distance metric and propose a unified one accordingly, which is applicable to ordinal-and-nominal-attribute categorical data. Compared with the existing metrics proposed for categorical data, the proposed metric is simple to use and nonparametric. More importantly, it reasonably exploits the underlying order information of ordinal attributes and statistical information of nominal attributes for distance measurement. Extensive experiments show that the proposed metric outperforms the existing counterparts on both the real and benchmark data sets. Yiqun Zhang 0006, Yiu-Ming Cheung, Kay Chen Tan |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2019 | Fast and Accurate Hierarchical Clustering Based on Growing Multilayer Topology TrainingabstractHierarchical clustering has been extensively applied for data analysis and knowledge discovery. However, the scalability of hierarchical clustering methods is generally limited due to their time complexity of O(n2), where n is the size of the input data. To address this issue, we present a fast and accurate hierarchical clustering algorithm based on topology training. Specifically, a trained multilayer topological structure that fits the spatial distribution of the data is utilized to accelerate the similarity measurement, which dominates the computational cost in hierarchical clustering. Moreover, the topological structure also guides the merging steps in hierarchical clustering to form a meaningful and accurate clustering result. In addition, an incremental version of the proposed algorithm is further designed so that the proposed approach is applicable to the streaming data as well. Promising experimental results on various data sets demonstrate the efficiency and effectiveness of the proposed algorithms. Yiu-Ming Cheung, Yiqun Zhang 0006 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2018 | Exploiting Order Information Embedded in Ordered Categories for Ordinal Data Clustering
Yiqun Zhang 0006, Yiu-Ming Cheung |
ISMIS | 1 |
| 2016 | Quality preserved data summarization for fast hierarchical clusteringabstractTraditional hierarchical clustering (HC) methods are not scalable with the size of databases. To address this issue, a series of summarization techniques, i.e. data bubbles (DB) and its improved versions, have been proposed to compress very large databases into representative seed points suitable for subsequent hierarchy construction. However, DB and its variants have two common drawbacks: (1) their performance is sensitive to the compression rate, and (2) their performance is sensitive to the initialization, i.e. the number and location of initialized seed points. This paper therefore proposes a new data summarization scheme, which is efficient and robust against the compression rate and initialization. In the proposed scheme, seed points are not only randomly initialized, but also trained to make them representative. After the training, a link strength network is constructed to achieve accurate hierarchy structure construction. Experiments demonstrate that the proposed method can produce high quality hierarchy structure with very high compression rate. Yiqun Zhang 0006, Yiu-Ming Cheung, Yang Liu 0007 |
IJCNN | 1 |