EDBT 2026 Demo / reviewers in the wild / expert
Junming Shao
dblp:93/7664
· DBLP profile ↗
53ranked-venue papers in the field
15as first author
23since 2021 · last 2026
—ORCID · conflict
Domains — venue-derived; a paper can count in several
Data Mining & Knowledge Discovery · 22 (10 first)Knowledge Engineering, Semantic Web & Information Systems · 19 (1 first)Database Systems & Data Management · 7 (4 first)Information Retrieval & Web Search · 4Other / Interdisciplinary · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Community-Level Personalized Recommendation by Exploiting Evolving User-Item Micro-Clusters
Jinxia Guo, Qirui Hao, Zhongjing Yu, Qinli Yang, Junming Shao |
ICDE | 6 |
| 2026 | A Unified Framework for Rule Learning: Integrating Commonsense Knowledge from LLMs with Structured Knowledge from Knowledge Graphs
Qirui Hao, Kewei Cheng, Tongze Zhang, Hongyuan Liu 0006, Junming Shao, Carl Yang 0001 |
WWW | 5 |
| 2026 | GAN-TAD: A graph-based adversarial network for time series anomaly detection
Tadiyos Hailemichael Mamo, Junming Shao, Cobbinah Bernard Mawuli |
Inf. Sci. | 2 |
| 2026 | Exploiting reliable evolving micro-clusters for robust semi-supervised learning on data streams
Zhonglin Wu, Jinxia Guo, Wei Han 0009, Qinli Yang, Junming Shao |
Inf. Sci. | 7 |
| 2024 | A reliable adaptive prototype-based learning for evolving data streams with limited labels
Cobbinah Bernard Mawuli, Qinli Yang, Junming Shao |
Inf. Process. Manag. | 5 |
| 2024 | Synchronization-based semi-supervised data streams classification with label evolution and extreme verification delay
Qinli Yang, Junming Shao, Cobbinah Bernard Mawuli, Waqar Ali 0001 |
Inf. Sci. | 3 |
| 2024 | Robust graph embedding via Attack-aid Graph Denoising
Zhili Qin, Zhongjing Yu, Qinli Yang, Junming Shao |
Inf. Sci. | 5 |
| 2024 | Unsupervised graph denoising via feature-driven matrix factorization
Zhili Qin, Zejun Sun, Qinli Yang, Junming Shao |
Inf. Sci. | 5 |
| 2024 | Learning evolving prototypes for imbalanced data stream classification with limited labels
Zhonglin Wu, Jingxia Guo, Qinli Yang, Junming Shao |
Inf. Sci. | 5 |
| 2023 | Learning multiple gaussian prototypes for open-set recognition
Jiaming Liu 0002, Wei Han 0009, Zhili Qin, Yulu Fan, Junming Shao |
Inf. Sci. | 6 |
| 2023 | Semi-supervised federated learning on evolving data streams
Cobbinah Bernard Mawuli, Jay Kumar, Ebenezer Nanor, Shangxuan Fu, Liangxu Pan, Qinli Yang, Junming Shao |
Inf. Sci. | 8 |
| 2023 | FedSULP: A communication-efficient federated learning framework with selective updating and loss penalization
Ebenezer Nanor, Cobbinah Bernard Mawuli, Qinli Yang, Junming Shao, Christiana Kobiah |
Inf. Sci. | 4 |
| 2022 | Learning Evolving Concepts with Online Class Posterior Probability
Junming Shao, Jianyun Lu, Zhili Qin, Qiming Wangyang, Qinli Yang |
DASFAA (2) | 1 |
| 2022 | Visualization in Data Science VDS @ KDD 2022abstractData science is the practice of deriving insight from data, enabled by modeling, computational methods, interactive visual analysis, and domain-driven problem solving. Data science draws from methodology developed in such fields as applied mathematics, statistics, machine learning, data mining, data management, visualization, and HCI. It drives discoveries in business, economy, biology, medicine, environmental science, the physical sciences, the humanities and social sciences, and beyond. Machine learning and data mining and visualization are integral parts of data science, and essential to enable sophisticated analysis of data. Nevertheless, both research areas are currently still rather separated and investigated by different communities rather independently. The goal of this workshop is to bring researchers from both communities together in order to discuss common interests, to talk about practical issues in application-related projects, and to identify open research problems. This summary gives a brief overview of the ACM KDD Workshop on Visualization in Data Science (VDS at ACM KDD and IEEE VIS), which will take place virtually on Aug 14-18, 2022 (Held in conjunction with KDD'22). The workshop website is available at http://www.visualdatascience.org/2022/ Claudia Plant, Nina C. Hubig, Junming Shao, Alvitta Ottley, Liang Gou, Torsten Möller, Adam Perer, Alexander Lex, Anamaria Crisan |
KDD | 3 |
| 2022 | Community detection in subspace of attribute
Zhongjing Yu, Qinli Yang, Junming Shao |
Inf. Sci. | 4 |
| 2022 | Multi-instance attention network for few-shot learning
Zhili Qin, Cobbinah Bernard Mawuli, Wei Han 0009, Rui Zhang 0070, Qinli Yang, Junming Shao |
Inf. Sci. | 7 |
| 2021 | Flexible, Robust, Scalable Semi-supervised Learning via Reliability PropagationabstractSemi-supervised learning aims to generate a model with a better performance using plenty of unlabeled data. However, most existing methods treat unlabeled data equally without considering whether it is safe or not, which may lead to the degradation of prediction performance. In this paper, towards reliable semi-supervised learning, we propose a data-driven algorithm, called Reliability Propagation (RP), to learn the reliability of each unlabeled instance. The basic idea is to take local label regularity as a prior, and then perform reliability propagation on an adaptive graph. As a result, the most reliable unlabeled instances could be selected to construct a safer classifier. Beyond, the distributed RP algorithm is introduced to scale up to large volumes of data. In contrast to existing approaches, RP exploits the structural information and shed light on the soft instance selection for unlabeled data in a classifier-independent way. Experiments on both synthetic and real-world data have demonstrated that RP allows extracting most reliable unlabeled instances and supports a gained prediction performance compared to other algorithms. Liangxu Pan, Qinli Yang, Junming Shao |
ICDM | 5 |
| 2021 | A general framework for mining concept-drifting data streams with evolvable featuresabstractMining feature evolvable streams has gained increasing attention in recent years. However, most existing approaches are designed for stationary data streams (i.e., data streams without concept drifts) and often work with high time complexity due to the time-consuming optimization procedure. The two deficiencies thus largely limit its applications to real-world data stream scenarios. In this paper, we consider a more difficult but practical streaming setting: a data stream with both concept drifts and evolvable features. To this end, we propose a general framework for mining concept-drifting data streams with evolvable features, called FEMC, based on an efficient Feature Evolvable streaming learning and dynamic Micro-Clusters maintenance. Specifically, we derive a closed-form solution to preserve the information in vanished features by learning a weight vector on survival features. The evolving concepts, are further learnt by dynamically maintaining a set of micro-clusters with varying feature space on-the-fly. Empirical results on real-world data sets have demonstrated the benefits of the proposed framework on both clustering and classification tasks by comparing with state-of-the-art algorithms. Jiaqi Peng, Jinxia Guo, Qinli Yang, Jianyun Lu, Junming Shao |
ICDM | 5 |
| 2021 | VDS'21: Visualization in Data ScienceabstractData science is the practice of deriving insight from data, enabled by modeling, computational methods, interactive visual analysis, and domain-driven problem solving. Data science draws from methodology developed in such fields as applied mathematics, statistics, machine learning, data mining, data management, visualization, and HCI. It drives discoveries in business, economy, biology, medicine, environmental science, the physical sciences, the humanities and social sciences, and beyond. Machine learning and data mining and visualization are integral parts of data science, and essential to enable sophisticated analysis of data. Nevertheless, both research areas are currently still rather separated and investigated by different communities rather independently. The goal of this workshop is to bring researchers from both communities together in order to discuss common interests, to talk about practical issues in application-related projects, and to identify open research problems. This summary gives a brief overview of the ACM KDD Workshop on Visualization in Data Science (VDS at ACM KDD and IEEE VIS), which will take place virtually on Aug 14-18, 2021 (Held in conjunction with KDD'21). The workshop website is available at: http://www.visualdatascience.org/2021/ Claudia Plant, Alvitta Ottley, Liang Gou, Torsten Möller, Adam Perer, Alexander Lex, Junming Shao |
KDD | 7 |
| 2021 | Learning Dynamic User Behavior Based on Error-driven Event RepresentationabstractUnderstanding the evolution of large graphs over time is of significant importance in user behavior understanding and prediction. Modeling user behavior with temporal networks has gained increasing attention in recent years since it allows capturing users’ dynamic preferences and predicting their next actions. Recently, some approaches have been proposed to model user behavior. However, these methods suffer from two problems: they work on static data, which ignores the dynamic evolution, or they model the whole behavior sequences directly by recurrent neural networks and thus suffer from noisy information. To tackle these problems, we propose a dynamic user behavior learning algorithm called LDBR. It views user behaviors as a set of dynamic events and uses recent event embedding to predict future user behavior and infer the current semantic labels. Specifically, we propose a new strategy to automatically learn a good event embedding in behavior sequence by introducing a smooth sampling strategy and minimizing the temporal link prediction error. Honglian Wang, Peiyan Li 0002, Wujun Tao, Bailin Feng, Junming Shao |
WWW | 5 |
| 2021 | Modular neural network via exploring category hierarchy
Wei Han 0009, Changgang Zheng, Rui Zhang 0070, Jinxia Guo, Qinli Yang, Junming Shao |
Inf. Sci. | 6 |
| 2021 | Towards real-time demand-aware sequential POI recommendation
Honglian Wang, Peiyan Li 0002, Junming Shao |
Inf. Sci. | 4 |
| 2021 | Data stream classification with novel class detection: a review, comparison and challenges
Junming Shao, Jay Kumar, Cobbinah Bernard Mawuli, S. M. Hasan Mahmud, Qinli Yang |
Knowl. Inf. Syst. | 2 |
| 2020 | Exploiting Inconsistency Problem in Multi-label Classification via Metric LearningabstractMulti-label classification problem has gained growing attention in recent years due to its diverse applications to real-world problems such as image annotation and query suggestions. However, traditional multi-label classification methods tend to fail due to the inconsistency between input and output space, where similar instances in the feature space may have distinct semantic labels in the output space. To eliminate the inconsistency problem, in this paper, we propose a supervised metric learning approach for multi-label classification, called MLMLI, which attempts to learn a similarity metric for multi-label data. The basic idea is to incorporate label similarity in output space as weak supervision to assign higher similarity to the pairs of instances with more similar labels. To this end, a weighted triple loss, and a step-specified coordinate descent method are employed. Different from traditional dimensionality reduction approaches, MLMLI is independent of any prior information of data, and thus enjoys a high capacity of generalization. Moreover, the metric learned by MLMLI offers a new venue for feature learning. Experiments on real-world datasets have further demonstrated the effectiveness of MLMLI and show its superiority over many state-of-the-art algorithms. Peiyan Li 0002, Zhili Qin, Honglian Wang, Qinli Yang, Junming Shao |
ICDM | 5 |
| 2020 | Community Detection with Local Metric LearningabstractCommunity detection in attributed networks has gained growing attention in recent years due to the booming of network data with both topological structure and attributes of nodes. To date, numerous algorithms have been proposed to leverage both kinds of information to yield high-quality communities based on homophily assumption (i.e., nodes are likely to link with those who share similar attributes). However, these approaches tend to focus on consistent information only and fail to consider the heterogeneity between topology and attributes. In light of the problem, we propose a new algorithm called CDLM (Community Detection via Local Metric learning) for attributed networks. The key point is to combine topological structure and node attributes in local metric space and perform community detection and local metric learning iteratively. With such a strategy, the learned local distance measures will benefit the performance of community detection, and in turn, the identification of intrinsic community structure helps to eliminate the negative effects of noisy edges in learning local metrics. Notably, homogeneity and heterogeneity between topological structure and node attributes are simultaneously considered to boost the performance of community detection. Experimental results on both synthetic and real-world networks have demonstrated the effectiveness of the proposed CDLM algorithm. Peiyan Li 0002, Honglian Wang, Jianyun Lu, Qinli Yang, Junming Shao |
ICDM | 5 |
| 2020 | Community Attention Network for Semi-supervised Node ClassificationabstractGraph neural networks (GNNs) have achieved great success for semi-supervised node classification by embedding node representation into a low-dimensional space. However, existing approaches usually ignore one intrinsic property of graphs: community structure, where the formation of distinct communities in graphs is often driven by different subset of attributes. In this paper, we introduce a new method, called Community Attention Network (CAT), aiming to extract community-specific features and then enhance node embeddings for classification. To learn such community-specific information, we design a new loss function to ensure the nodes in the same community should share similar attributes (i.e., low covariance), and any unlabelled node should belong to only one class with high probability (i.e., low community distribution entropy) in a community attention network. Extensive experimental results demonstrate the effectiveness of CAT and its advantages over many state-of-the-art approaches. To further illustrate the benefits of CAT to capture the community information, a case study is given and discussed. Zhongjing Yu, Christian Böhm 0001, Junming Shao |
ICDM | 5 |
| 2020 | Learning evolving user's behaviors on location-based social networks
Ruizhi Wu, Guangchun Luo, Junming Shao, Chang-Tien Lu |
GeoInformatica | 4 |
| 2020 | Attributed graph clustering with subspace stochastic block model
Zhongjing Yu, Qinli Yang, Junming Shao |
Inf. Sci. | 4 |
| 2020 | Exploiting evolving micro-clusters for data stream classification with emerging class detection
Junming Shao |
Inf. Sci. | 2 |
| 2020 | Online reliable semi-supervised learning on evolving data streams
Junming Shao, Jay Kumar, Waqar Ali 0001, Jiaming Liu 0002 |
Inf. Sci. | 2 |
| 2020 | Semantic trajectory representation and retrieval via hierarchical embedding
Chongming Gao, Zhong Zhang 0004, Hongzhi Yin, Qinli Yang, Junming Shao |
Inf. Sci. | 6 |
| 2020 | Structured subspace embedding on attributed networks
Zhongjing Yu, Zhong Zhang 0004, Junming Shao |
Inf. Sci. | 4 |
| 2019 | Towards Robust Arbitrarily Oriented Subspace Clustering
Zhong Zhang 0004, Chongming Gao, Chongzhi Liu, Qinli Yang, Junming Shao |
DASFAA (1) | 5 |
| 2019 | Online Budgeted Least Squares with Unlabeled DataabstractThe scarcity of labeled data in real streaming environments has boosted the study of online semi-supervised learning (SSL). However, existing online SSL models often rely on some specific assumptions (e.g., manifold assumption) and need to maintain some extra constraints (e.g., the Laplacian matrix) on the fly, which is usually time and resource consuming. In this paper, we propose an efficient and effective online semi-supervised learning approach via Budgeted Least Square (BLS). Specifically, we first derive both closed-form transductive and inductive solutions for kernel least squares classification in the semi-supervised setting. Then, together with online kernel learning, BLS allows a concise online update. Besides, the theoretical regret bound of BLS is analysed, and empirical experiments on both static and streaming data further demonstrate its superiority over state-of-the-art algorithms. Peiyan Li 0002, Chongming Gao, Qinli Yang, Junming Shao |
ICDM | 5 |
| 2019 | Synchronization-based clustering on evolving data stream
Junming Shao, Lianli Gao, Qinli Yang, Claudia Plant, Ira Assent |
Inf. Sci. | 1 |
| 2019 | Preference modeling by exploiting latent components of ratings
Wei Zeng 0013, Junming Shao, Ge Fan |
Knowl. Inf. Syst. | 3 |
| 2018 | Graph Clustering with Local Density-Cut
Junming Shao, Qinli Yang, Zhong Zhang 0004, Jinhu Liu, Stefan Kramer 0001 |
DASFAA (1) | 1 |
| 2018 | Multi-view Discriminative Learning via Joint Non-negative Matrix Factorization
Zhong Zhang 0004, Zhili Qin, Peiyan Li 0002, Qinli Yang, Junming Shao |
DASFAA (2) | 5 |
| 2018 | Robust Prototype-Based Learning on Data StreamsabstractIn this paper, we propose a prototype-based classification model for evolving data streams, called SyncStream, which allows dynamically modeling time-changing concepts, making predictions in a local fashion. Instead of learning a single model on a fixed or adaptive sliding window of historical data or ensemble learning a set of weighted base classifiers, SyncStream captures evolving concepts by dynamically maintaining a set of prototypes in a proposed P-Tree, which are obtained based on the error-driven representativeness learning and synchronization-inspired constrained clustering. To identify abrupt concept drifts in data streams, PCA and statistical analysis based heuristic approaches have been introduced. To further learn the associations among distributed data streams, the extended P-Tree structure and KNN-style strategy are introduced. We demonstrate that our new data stream classification approach has several attractive benefits: (a) SyncStream is capable of dynamically modeling the evolving concepts from even a small set of prototypes. (b) Owing to synchronization-based constrained clustering and P-Tree, SyncStream supports efficient and effective data representation and maintenance. (c) SyncStream is also tolerant of inappropriate or noisy examples via error-driven representativeness learning. (d) SyncStream allows learning relationship among distributed data streams at the instance level. The experimental results indicate its efficiency and effectiveness. Junming Shao, Qinli Yang, Guangchun Luo |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2017 | Synchronization-Inspired Co-Clustering and Its Application to Gene Expression DataabstractIn this paper, we propose a new synchronization-inspired co-clustering algorithm by dynamic simulation, called CoSync, which aims to discover biologically relevant subgroups embedding in a given gene expression data matrix. The basic idea is to view a gene expression data matrix as a dynamical system, and the weighted two-sided interactions are imposed on each element of the matrix from both aspects of genes and conditions, resulting in the values of all element in a co-cluster synchronizing together. Experiments show that our algorithm allows uncovering high-quality co-clusterings embedded in gene expression data sets and has its superiority over many state-of-the-art algorithms. Junming Shao, Chongming Gao, Wei Zeng 0013, Jingkuan Song, Qinli Yang |
ICDM | 1 |
| 2017 | Exploring Common and Distinct Structural Connectivity Patterns Between Schizophrenia and Major Depression via Cluster-Driven Nonnegative Matrix FactorizationabstractIn this paper, we introduce a novel method to discover common and distinct structural connectivity patterns between SZP and MDD via a Cluster-Driven Nonnegative Matrix Factorization (called CD-NMF). Specifically, CD-NMF is applied to decompose the joint structural connectivity map into common and distinct parts, and each part is further factorized into two sub-matrices (i.e. common/distinct basis matrix and common/distinct encoding matrix) correspondingly. By imposing the clustering constraints on common and distinct encoding matrices, the discriminative patterns as well as the common patterns between the two disorders are extracted simultaneously. Experimental results demonstrate that CD-NMF allows finding the common and distinct structural patterns effectively. More importantly, the derived distinct patterns, show powerful ability to discriminate the patients of schizophrenia and major depressive disorder. Junming Shao, Zhongjing Yu, Peiyan Li 0002, Wei Han 0009, Christian Sorg, Qinli Yang |
ICDM | 1 |
| 2017 | Synchronization-based scalable subspace clustering of high-dimensional data
Junming Shao, Xinzuo Wang, Qinli Yang, Claudia Plant, Christian Böhm 0001 |
Knowl. Inf. Syst. | 1 |
| 2016 | Reliable Semi-supervised LearningabstractIn this paper, we propose a Reliable Semi-Supervised Learning framework, called ReSSL, for both static and streaming data. Instead of relaxing different assumptions, we do model the reliability of cluster assumption, quantify the distinct importance of clusters (or evolving micro-clusters on data streams), and integrate the cluster-level information and labeled data for prediction with a lazy learning framework. Extensive experiments demonstrate that our method has good performance compared to state-of-the-art algorithms on data sets in both static and real streaming environments. Junming Shao, Qinli Yang, Guangchun Luo |
ICDM | 1 |
| 2016 | Scalable Clustering by Iterative Partitioning and Point Attractor RepresentationabstractClustering very large datasets while preserving cluster quality remains a challenging data-mining task to date. In this paper, we propose an effective scalable clustering algorithm for large datasets that builds upon the concept of synchronization. Inherited from the powerful concept of synchronization, the proposed algorithm, CIPA (Clustering by Iterative Partitioning and Point Attractor Representations), is capable of handling very large datasets by iteratively partitioning them into thousands of subsets and clustering each subset separately. Using dynamic clustering by synchronization, each subset is then represented by a set of point attractors and outliers. Finally, CIPA identifies the cluster structure of the original dataset by clustering the newly generated dataset consisting of points attractors and outliers from all subsets. We demonstrate that our new scalable clustering approach has several attractive benefits: (a) CIPA faithfully captures the cluster structure of the original data by performing clustering on each separate data iteratively instead of using any sampling or statistical summarization technique. (b) It allows clustering very large datasets efficiently with high cluster quality. (c) CIPA is parallelizable and also suitable for distributed data. Extensive experiments demonstrate the effectiveness and efficiency of our approach. Junming Shao, Qinli Yang, Hoang-Vu Dang, Bertil Schmidt, Stefan Kramer 0001 |
ACM Trans. Knowl. Discov. Data | 1 |
| 2015 | Community Detection based on Distance DynamicsabstractIn this paper, we introduce a new community detection algorithm, called Attractor, which automatically spots communities in a network by examining the changes of "distances" among nodes (i.e. distance dynamics). The fundamental idea is to envision the target network as an adaptive dynamical system, where each node interacts with its neighbors. The interaction will change the distances among nodes, while the distances will affect the interactions. Such interplay eventually leads to a steady distribution of distances, where the nodes sharing the same community move together and the nodes in different communities keep far away from each other. Building upon the distance dynamics, Attractor has several remarkable advantages: (a) It provides an intuitive way to analyze the community structure of a network, and more importantly, faithfully captures the natural communities (with high quality). (b) Attractor allows detecting communities on large-scale networks due to its low time complexity (O(|E|)). (c) Attractor is capable of discovering communities of arbitrary size, and thus small-size communities or anomalies, usually existing in real-world networks, can be well pinpointed. Extensive experiments show that our algorithm allows the effective and efficient community detection and has good performance compared to state-of-the-art algorithms. Junming Shao, Zhichao Han 0003, Qinli Yang, Tao Zhou 0001 |
KDD | 1 |
| 2015 | Zero-shot Image Categorization by Image Correlation ExplorationabstractThe problem of image categorization from zero or only a few training examples, called zero-shot learning, occurs frequently, but it has hardly been studied in computer vision research. To tackle this problem, mid-level semantic attributes are introduced to identify image categories. For example, one can construct a classifier for the giant panda category by enumerating its attributes (e.g., black, white and four-footed) even without providing giant panda training images. Recently, several studies have investigated to learn attribute classifiers, based on which new classes can be detected. However, an often-encountered problem is the limited number of training data due to the time-consuming manual annotation of the attributes. Also, using single feature is hard to detect some attributes, e.g., the HSV feature is not robust enough to predict 'tusk' or 'flies' attributes. In this paper, we propose a unified semi-supervised learning (SSL) framework that learns the attribute classifiers by utilizing multiple feature and exploring the correlations between images. Specifically, we learn an optimal graph which embeds the relationships among the data points more accurately. Then, this graph is used to generate a geometrical regularizers for a semi-supervised learning model to learn the attribute classifier by utilizing both labeled and unlabeled images. Afterward, new classes can be detected based on their attribute representation. The use of SSL can boost the performances of attribute classifiers with very few training examples, and the adoption of multiple features makes the attribute prediction more robust. Experimental results on a series of real benchmark data sets suggest that semi-supervised learning do enhance the performances of attribute prediction and zero-shot categorization, compared with state-of-the-art methods. Lianli Gao, Jingkuan Song, Junming Shao, Xiaofeng Zhu 0001, Heng Tao Shen |
ICMR | 3 |
| 2014 | Prototype-based learning on concept-drifting data streamsabstractData stream mining has gained growing attentions due to its wide emerging applications such as target marketing, email filtering and network intrusion detection. In this paper, we propose a prototype-based classification model for evolving data streams, called SyncStream, which dynamically models time-changing concepts and makes predictions in a local fashion. Instead of learning a single model on a sliding window or ensemble learning, SyncStream captures evolving concepts by dynamically maintaining a set of prototypes in a new data structure called the P-tree. The prototypes are obtained by error-driven representativeness learning and synchronization-inspired constrained clustering. To identify abrupt concept drift in data streams, PCA and statistics based heuristic approaches are employed. SyncStream has several attractive benefits: (a) It is capable of dynamically modeling evolving concepts from even a small set of prototypes and is robust against noisy examples. (b) Owing to synchronization-based constrained clustering and the P-Tree, it supports an efficient and effective data representation and maintenance. (c) Gradual and abrupt concept drift can be effectively detected. Empirical results shows that our method achieves good predictive performance compared to state-of-the-art algorithms and that it requires much less time than another instance-based stream mining algorithm. Junming Shao, Zahra Ahmadi, Stefan Kramer 0001 |
KDD | 1 |
| 2013 | Robust Synchronization-Based Graph Clustering
Junming Shao, Xiao He 0002, Qinli Yang, Claudia Plant, Christian Böhm 0001 |
PAKDD (1) | 1 |
| 2013 | Synchronization-Inspired Partitioning and Hierarchical ClusteringabstractSynchronization is a powerful and inherently hierarchical concept regulating a large variety of complex processes ranging from the metabolism in a cell to opinion formation in a group of individuals. Synchronization phenomena in nature have been widely investigated and models concisely describing the dynamical synchronization process have been proposed, e.g., the well-known Extensive Kuramoto Model. We explore the potential of the Extensive Kuramoto Model for data clustering. We regard each data object as a phase oscillator and simulate the dynamical behavior of the objects over time. By interaction with similar objects, the phase of an object gradually aligns with its neighborhood, resulting in a nonlinear object movement naturally driven by the local cluster structure. We demonstrate that our framework has several attractive benefits: 1) It is suitable to detect clusters of arbitrary number, shape, and data distribution, even in difficult settings with noise points and outliers. 2) Combined with the Minimum Description Length (MDL) principle, it allows partitioning and hierarchical clustering without requiring any input parameters which are difficult to estimate. 3) Synchronization faithfully captures the natural hierarchical cluster structure of the data and MDL suggests meaningful levels of abstraction. Extensive experiments demonstrate the effectiveness and efficiency of our approach. Junming Shao, Xiao He 0002, Christian Böhm 0001, Qinli Yang, Claudia Plant |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2011 | Detection of Arbitrarily Oriented Synchronized Clusters in High-Dimensional DataabstractHow to address the challenges of the "curse of dimensionality" in clustering? Clustering is a powerful data mining technique for structuring and organizing vast amounts of data. However, the high-dimensional data space is usually very sparse and meaningful clusters can only be found in lower dimensional subspaces. In many applications the subspaces hosting the clusters provide valuable information for interpreting the major patterns in the data. Detection of subspace clusters is challenging since usually many of the attributes are noisy, some attributes may exhibit correlations among each other and only few of the attributes truly contribute to the cluster structure. In this paper, we propose ORSC (Arbitrarily ORiented Synchronized Clusters), a novel effective and efficient method to subspace clustering inspired by synchronization. Synchronization is a basic phenomenon prevalent in nature, capable of controlling even highly complex processes such as opinion formation in a group. Control of complex processes is achieved by simple operations based on interactions between objects. Relying on the interaction model for synchronization, our approach ORSC (1) naturally detects correlation clusters in arbitrarily oriented subspaces, including (2) arbitrarily shaped non-linear correlation clusters. Our approach is (3) robust against noise points and outliers. In contrast to previous methods, ORSC is (4) easy to parameterize, since there is no need to specify the subspace dimensionality and all interesting subspace clusters can be detected. Finally, (5) ORSC outperforms most comparison methods in terms of runtime efficiency and is highly scalable to large and high-dimensional data sets. Junming Shao, Claudia Plant, Qinli Yang, Christian Böhm 0001 |
ICDM | 1 |
| 2011 | Weighted Graph Compression for Parameter-free Clustering With PaCCoabstractObject similarities are now more and more characterized by connectivity information available in form of network or graph data.Complex graph data arises in various fields like e-commerce, social networks, high throughput biological analysis etc.The generated interaction information for objects is often not simply binary but rather associated with interaction strength which are in turn represented as edge weights in graphs.The identification of groups of highly connected nodes is an important task and results in valuable knowledge of the data set as a whole.Many popular clustering techniques are designed for vector or unweighted graph data, and can thus not be directly applied for weighted graphs.In this paper, we propose a novel clustering algorithm for weighted graphs, called PaCCo (Parameter-free C lustering by Coding costs), which is based on the Minimum Description Length (MDL) principle in combination with a bisecting k-Means strategy.MDL relates the clustering problem to the problem of data compression: A good cluster structure on graphs enables strong graph compression.The compression efficiency depends on the underlying edges which constitute the graph connectivity.The compression rate serves as similarity or distance metric for nodes.The MDL principle ensures that our algorithm is parameter free (automatically finds the number of clusters) and avoids restrictive assumptions that no information on the data is required.We systematically evaluate our clustering approach PaCCo on synthetic as well as on real data to demonstrate the superiority of our developed algorithm over existing approaches. Nikola S. Müller, Katrin Haegler, Junming Shao, Claudia Plant, Christian Böhm 0001 |
SDM | 3 |
| 2010 | Clustering by synchronizationabstractSynchronization is a powerful basic concept in nature regulating a large variety of complex processes ranging from the metabolism in the cell to social behavior in groups of individuals. Therefore, synchronization phenomena have been extensively studied and models robustly capturing the dynamical synchronization process have been proposed, e.g. the Extensive Kuramoto Model. Inspired by the powerful concept of synchronization, we propose Sync, a novel approach to clustering. The basic idea is to view each data object as a phase oscillator and simulate the interaction behavior of the objects over time. As time evolves, similar objects naturally synchronize together and form distinct clusters. Inherited from synchronization, Sync has several desirable properties: The clusters revealed by dynamic synchronization truly reflect the intrinsic structure of the data set, Sync does not rely on any distribution assumption and allows detecting clusters of arbitrary number, shape and size. Moreover, the concept of synchronization allows natural outlier handling, since outliers do not synchronize with cluster objects. For fully automatic clustering, we propose to combine Sync with the Minimum Description Length principle. Extensive experiments on synthetic and real world data demonstrate the effectiveness and efficiency of our approach. Christian Böhm 0001, Claudia Plant, Junming Shao, Qinli Yang |
KDD | 3 |
| 2010 | Synchronization Based Outlier Detection
Junming Shao, Christian Böhm 0001, Qinli Yang, Claudia Plant |
ECML/PKDD (3) | 1 |