EDBT 2026 Demo / reviewers in the wild / expert
Xiaopeng Luo
dblp:139/0320
· DBLP profile ↗
15ranked-venue papers
2as first author
14since 2021 · last 2026
0000-0002-9806-8892ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Mask the Redundancy: Evolving Masking Representation Learning for Multivariate Time-Series ClusteringabstractMultivariate Time-Series (MTS) clustering discovers intrinsic grouping patterns of temporal data samples. Although time-series provide rich discriminative information, they also contain substantial redundancy, such as steady-state machine operation records and zero-output periods of solar power generation. Such redundancy diminishes the attention given to discriminative timestamps in representation learning, thus leading to performance bottlenecks in MTS clustering. Masking has been widely adopted to enhance the MTS representation, where temporal reconstruction tasks are designed to capture critical information from MTS. However, most existing masking strategies appear to be standalone preprocessing steps, isolated from the learning process, which hinders dynamic adaptation to the importance of clustering-critical timestamps. Accordingly, this paper proposes the Evolving-masked MTS Clustering (EMTC) method, whose model architecture comprises Importance-aware Variate-wise Masking (IVM) and Multi-Endogenous Views (MEV) generation modules. IVM adaptively guides the model in learning more discriminative representations for clustering, while the reconstruction and cluster-guided contrastive learning pathways enhance and connect the representation learning to clustering tasks. Extensive experiments on 15 benchmark datasets demonstrate the superiority of EMTC over eight SOTA methods, where the EMTC achieves an average improvement of 4.85% in F1-Score over the strongest baselines. Zexi Tan, Xiaopeng Luo, Yunlin Liu, Yiqun Zhang 0006 |
AAAI | 2 |
| 2026 | Hierarchical Reference Sets for Robust Unsupervised Detection of Scattered and Clustered Outliersabstractreal-world IoT data analysis tasks, such as clustering and anomaly event detection, are unsupervised and highly susceptible to the presence of outliers. In addition to sporadic scattered outliers caused by factors such as faulty sensor readings, IoT systems often exhibit clustered outliers. These occur when multiple devices or nodes produce similar anomalous measurements, for instance, owing to localized interference, emerging security threats, or regional false alarms, forming micro-clusters. These clustered outliers can be easily mistaken for normal behavior because of their relatively high local density, thereby obscuring the detection of both scattered and contextual anomalies. To address this, we propose a novel outlier detection paradigm that leverages the natural neighboring relationships using graph structures. This facilitates multi-perspective anomaly evaluation by incorporating reference sets at both local and global scales derived from the graph. Our approach enables the effective recognition of scattered outliers without interference from clustered anomalies, whereas the graph structure simultaneously helps reflect and isolate clustered outlier groups. Extensive experiments, including comparative performance analysis, ablation studies, validation on downstream clustering tasks, and evaluation of hyperparameter sensitivity, demonstrate the efficacy of the proposed method. The source code is available at https://github.com/gordonlok/DROD. Yiqun Zhang 0006, Zexi Tan, Xiaopeng Luo, Yunlin Liu |
IEEE Internet Things J. | 3 |
| 2025 | Robust Unsupervised Outlier Detection in Mixed Data Using Hierarchical Reference SetsabstractUnsupervised outlier detection in mixed-attribute data poses significant challenges in healthcare and network security domains where data combine numerical, nominal, and ordinal features. Existing methods struggle with two critical limitations: they typically handle only single-type data and fail to distinguish scattered outliers from clustered outliers-locally dense microclusters that are globally abnormal but mask each other due to internal consistency. This paper proposes HAOD (Heterogeneous Attribute-based Outlier Detection), a hierarchical framework that constructs Natural Neighbor Sets (NNS) for adaptive local structure modeling and organizes them into Natural Neighbor Graph Reference Sets (NGS) for global connectivity representation. A dependency-aware mixed-distance metric unifies heterogeneous attributes by quantifying inter-attribute correlations. HAOD integrates Local Isolation Score (LIS) and Subset Isolation Score (SIS) to comprehensively detect both anomaly types without the masking effect. Experiments on NSL-KDD, UNSW-NB15, and Thyroid Disease datasets show HAOD outperforms eight baseline methods across AUC, precision, and average precision metrics. Ablation studies confirm both components are essential. The method operates parameter-free with$\mathbf{O}\left(\mathbf{n}^{2} \mathbf{d}\right)$complexity, offering a robust solution for heterogeneous monitoring systems. Xiaopeng Luo, Zhuowei Wang 0001 |
BIBM | 1 |
| 2025 | MACL: Metric and Attribute Space Co-learning for Qualitative Data Clustering
Haoyi Xiao, Xinxi Chen, Xiaopeng Luo, Gengwen Huang |
ICIC (12) | 3 |
| 2025 | Learning Self-Growth Maps for Fast and Accurate Imbalanced Streaming Data ClusteringabstractStreaming data clustering is a popular research topic in data mining and machine learning. Since streaming data is usually analyzed in data chunks, it is more susceptible to encountering the dynamic cluster imbalance issue. That is, the imbalance ratio (IR) of clusters changes over time, which can easily lead to fluctuations in either the accuracy or the efficiency of streaming data clustering. Therefore, an accurate and efficient streaming data clustering approach is proposed to adapt to the drifting and imbalanced cluster distributions. We first design a self-growth map (SGM) that can automatically arrange neurons on demand according to local distribution, and thus achieve fast and incremental adaptation to the streaming distributions. Since SGM allocates an excess number of density-sensitive neurons to describe the global distribution, it can avoid missing small clusters among imbalanced distributions. We also propose a fast hierarchical merging (HM) strategy to combine the neurons that break up the relatively large clusters. It exploits the maintained SGM to quickly retrieve the intracluster distribution pairs for merging, which circumvents the most laborious global searching. It turns out that the proposed SGM can incrementally adapt to the distributions of new chunks, and the self-growth map-guided hierarchical merging for the imbalanced data clustering (SOHI) approach can quickly explore a true number of imbalanced clusters. Extensive experiments demonstrate that SOHI can efficiently and accurately explore cluster distributions for streaming data. Yiqun Zhang 0006, Sen Feng, Zexi Tan, Xiaopeng Luo, Yuzhu Ji, Yiu-Ming Cheung |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | MGOD: Multi-Granular Outlier Detection with Clustlier AnalysisabstractUnsupervised Outlier Detection (UOD) is crucial for the analysis of biomedical and health data with undesirable outliers. However, the complex distribution of real data often brings difficulties to UOD where the "masking effect", i.e., only a small number of densely distributed outliers (also called clustliers) can collectively mask themselves from being detected, is particularly challenging. Another difficulty derived from this is how to distinguish clustliers from small clusters. Therefore, we propose a novel Multi-Granular Outlier Detector (MGOD). It first partitions the dataset into subsets with natural neighbor topological relationships to circumvent the non-trivial neighbor range setting. Then it effectively detects both clustliers and isolated samples (also called scatliers) based on a newly designed anomaly score. The score comprehensively takes into account the density and connectivity of samples to reflect different extents and types of abnormality. It turns out that MGOD is accurate and highly interpretable. The performance of MGOD is also robust to the involved hyper-parameters, which are easy to set. Comprehensive evaluations have been conducted to compare seven counterparts on 15 datasets, most of which are biomedical datasets. The results of significance tests confirm the effectiveness and superiority of MGOD. The source code is opened at https://anonymous.4open.science/r/MGOD-C531. Qingsheng Chen, Mingjie Zhao 0003, Yuzhu Ji, Xiaopeng Luo, Yiqun Zhang 0006, Yue Zhang 0045 |
BIBM | 4 |
| 2024 | Efficient Topology-Driven Clustering for Imbalanced Streaming Biomedical Data AnalysisabstractClustering drifting data is common in the field of biomedical data analysis. Data chunks collected at different periods often exhibit clusters with significantly different sizes, and drifting distributions of clusters also appear frequently. We call such composite phenomenon imbalance-drifting, which can severely impact the accuracy and efficiency of cluster analysis. Therefore, we propose a topology-representation-based clustering paradigm, which first learns an informative global data representation in a self-organizing manner to obtain a map with nested representative data points. Then fast and accurate clustering is facilitated by quickly retrieving similar data points according to the topology. As the constructed Self-Organizing Map (SOM) is exploited for informative representation, micro partition, and quick merging, to achieve advanced clustering under imbalance-drifting, the proposed approach is thus called Tri-Squeezing SOM for Clustering (TSSC). It turns out that TSSC significantly reduces the time complexity for clustering an n-scale imbalance-streaming data without sacrificing accuracy. Moreover, TSSC can automatically determine the number of clusters k, and features interpretability and hyper-parameter robustness. Extensive results on both biomedical datasets and synthetic datasets verify the superiority of TSSC. Xiaopeng Luo, Yiqun Zhang 0006, Yuzhu Ji, Peng Liu 0045, Taoting Xiao |
BIBM | 1 |
| 2024 | Robust Categorical Data Clustering Guided by Multi-Granular Competitive LearningabstractData set composed of categorical features is very common in big data analysis tasks. Since categorical features are usually with a limited number of qualitative possible values, the nested granular cluster effect is prevalent in the implicit discrete distance space of categorical data. That is, data objects frequently overlap in space or subspace to form small compact clusters, and similar small clusters often form larger clusters. However, the distance space cannot be well-defined like the Euclidean distance due to the qualitative categorical data values, which brings great challenges to the cluster analysis of categorical data. In view of this, we design a Multi-Granular Competitive Penalization Learning (MGCPL) algorithm to allow potential clusters to interactively tune themselves and converge in stages with different numbers of naturally compact clusters. To leverage MGCPL, we also propose a Cluster Aggregation strategy based on MGCPL Encoding (CAME) to first encode the data objects according to the learned multi-granular distributions, and then perform final clustering on the embeddings. It turns out that the proposed MGCPL-guided Categorical Data Clustering (MCDC) approach is competent in automatically exploring the nested distribution of multi-granular clusters and highly robust to categorical data sets from various domains. Benefiting from its linear time complexity, MCDC is scalable to large-scale data sets and promising in pre-partitioning data sets or compute nodes for boosting distributed computing. Extensive experiments with statistical evidence demonstrate its superiority compared to state-of-the-art counterparts on various real public data sets. Shenghong Cai, Yiqun Zhang 0006, Xiaopeng Luo, Yiu-Ming Cheung, Hong Jia, Peng Liu 0045 |
ICDCS | 3 |
| 2024 | Towards Unbiased Minimal Cluster Analysis of Categorical-and-Numerical Attribute Data
Xiaopeng Luo, Qingsheng Chen, Yiqun Zhang 0006, Yiu-Ming Cheung |
ICPR (2) | 2 |
| 2023 | Selecting Heterogeneous Features Based on Unified Density-Guided Neighborhood Relation for Complex Biomedical Data AnalysisabstractBiomedical big data are usually high dimensional and collected in the form of a continuous influx of new features. Online Feature Selection (OFS) is a promising way to manage and analyze such data, as OFS circumvents the huge computation cost brought by simultaneously considering all the features, and can also dynamically maintain a distribution-fitting feature subset on the fly. However, almost all the OFS solutions are based on a naive premise that all features are of the same type, overlooking the fact that real biomedical data set usually consists of heterogeneous numerical and categorical features. This paper therefore proposes a new approach to Online Heterogeneous Feature Selection (OHFS), which dynamically maintains a feature subset that maximizes the number of neighborhood sets where all the objects within each neighborhood set are of the same class. To appropriately partition the objects into neighborhood sets, a density-guided relation is proposed, which adaptively forms non-overlapping neighborhood sets by detecting spatially compact objects. A unified density measure is also presented to avoid information loss in processing heterogeneous features. It turns out that the proposed approach features parameter- free, interpretability, and efficiency. It is capable of maintaining a concise feature subset while receiving any type of feature. Extensive experimental evaluations demonstrate its superiority. Lang Zhao, Yiqun Zhang 0006, Xiaopeng Luo, Yue Zhang 0045, Yiu-Ming Cheung, Kangshun Li |
BIBM | 3 |
| 2022 | Meta Distribution Alignment for Generalizable Person Re-IdentificationabstractDomain Generalizable (DG) person ReID is a challenging task which trains a model on source domains yet generalizes well on target domains. Existing methods use source domains to learn domain-invariant features, and assume those features are also irrelevant with target domains. However, they do not consider the target domain information which is unavailable in the training phrase of DG. To address this issue, we propose a novel Meta Distribution Alignment (MDA) method to enable them to share similar distribution in a test-time-training fashion. Specifically, since high-dimensional features are difficult to constrain with a known simple distribution, we first introduce an intermediate latent space constrained to a known prior distribution. The source domain data is mapped to this latent space and then reconstructed back. A meta-learning strategy is introduced to facilitate generalization and support fast adaption. To reduce their discrepancy, we further propose a test-time adaptive updating strategy based on the latent space which efficiently adapts model to unseen domains with a few samples. Extensive experimental results show that our model outperforms the state-of-the-art methods by up to 5.2% R-1 on average on the large-scale and 4.7% R-1 on the single-source domain generalization ReID benchmark. Source code is publicly available at https://github.com/haoni0812/MDA.git. Hao Ni 0002, Jingkuan Song, Xiaopeng Luo, Feng Zheng 0001, Wen Li 0001, Heng Tao Shen |
CVPR | 3 |
| 2022 | Heterogeneous Drift Learning: Classification of Mix-Attribute Data with Concept DriftsabstractAs many real data sets (e.g., social, financial, and medical data sets) are successively generated in evolution with the ever-changing environment, classification for data stream with concept drift attracts increasing attention in the fields of machine learning and data mining. However, to the best of our knowledge, existing works mainly consider the concept drift issue while ignoring another common characteristic of real data, i.e., existence of awkward heterogeneity caused by mixture of numerical and categorical attributes. It is worth noting that tackling both the concept drift and heterogeneity problems together is exponentially more challenging than dealing with only one of them. This paper, therefore, proposes an ensemble learning approach for the classification of numerical-and-categorical-attribute data (also called mixed data hereinafter) under concept drift. We first design a unified metric to appropriately address the heterogeneity of numerical and categorical attributes. Then a base classifier that can appropriately fuse the information provided by the heterogeneous attributes is formed accordingly. Furthermore, to make the classification adapt to the complex concept drifts demonstrated on the heterogeneous attributes, two types of base classifier ensembles are dynamically learned on the fly. Experimental results on various real mixed data sets with concept drifts demonstrate the efficacy of the proposed method. Lang Zhao, Yiqun Zhang 0006, Yuzhu Ji, An Zeng, Fangqing Gu, Xiaopeng Luo |
DSAA | 6 |
| 2022 | Global Convergence of Noisy Gradient DescentabstractNoise plays an important role in the gradient-based optimization methods, and a series of numerical experiments have demonstrated that adding gradient noise improves learning for neural networks. However, the mathematical interpretation of the noise remains a challenge. In this paper, we show that, the noise variation can be regarded as a smoothing factor, and we prove that, under certain conditions, a noisy gradient descent (NG) enjoys linear global convergence in expectation sense. We contribute to this problem by introducing an intermediate which connect the NG method to the smoothed function. On the one hand, this connection reveals that applying the NG method to a function is the same as applying the gradient method to the corresponding function smoothed by the noise; and on the other hand, it allows us to establish the convergence behavior of the NG in a global sense. Moreover, we also consider what conditions make the global minimizer of the smoothed function not far from the original global minimizer. Xuliang Qin, Xin Xu 0006, Xiaopeng Luo |
SMC | 3 |
| 2022 | Automatic Tagging by Leveraging Visual and Annotated Features in Social MediaabstractAutomatic image annotation is one of the research fields helping to extract the meaning of images, which aims at the production of a set of semantic annotations for an image to help better present the concept. Over the past few decades, researchers have developed many approaches for automatic image annotation. Nevertheless, previous studies have not fully accounted for visual features and annotated features. Therefore, it is still possible to achieve a better annotation performance by combining visual and annotated information. In this study, we aim to associate multiple semantic tags with a given image. In particular, we detect how to obtain the image annotation by utilizing visual and annotated information. To take advantage of visual information, we first designed a modified neural network method to acquire the features of the image content. In addition, to obtain the annotated features, we exploit an aggregated network embedding approach that consists of annotation embedding, social embedding, profile embedding, and semantic embedding. Finally, to produce an accurate image annotation, we integrate the two aforementioned methods, that is, combining the visual and annotated information, to build a unified cooperative training framework. The experimental results on three real-world datasets clarify that our presented method is superior to the currently popular image annotation approaches. Jinpeng Chen 0001, Pinguang Ying, Xiangling Fu, Xiaopeng Luo, Kaimin Wei |
IEEE Trans. Multim. | 4 |
| 2020 | Listening to the investors: A novel framework for online lending default prediction using deep learning neural networks
Xiangling Fu, Tianxiong Ouyang, Jinpeng Chen 0001, Xiaopeng Luo |
Inf. Process. Manag. | 4 |