EDBT 2026 Demo / reviewers in the wild / expert
Hongfu Liu 0001
dblp:32/9075-1
· DBLP profile ↗
69ranked-venue papers
19as first author
26since 2021 · last 2025
0000-0002-0821-8640ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 54 · 11 first-author · 24 since 2021Databases, data management, data science and information retrieval · 26 · 15 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 first-authorTheory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Outlier Gradient Analysis: Efficiently Identifying Detrimental Training Samples for Deep Learning ModelsabstractA core data-centric learning challenge is the identification of training samples that are detrimental to model performance. Influence functions serve as a prominent tool for this task and offer a robust framework for assessing training data influence on model predictions. Despite their widespread use, their high computational cost associated with calculating the inverse of the Hessian matrix pose constraints, particularly when analyzing large-sized deep models. In this paper, we establish a bridge between identifying detrimental training samples via influence functions and outlier gradient detection. This transformation not only presents a straightforward and Hessian-free formulation but also provides insights into the role of the gradient in sample impact. Through systematic empirical evaluations, we first validate the hypothesis of our proposed outlier gradient analysis approach on synthetic datasets. We then demonstrate its effectiveness in detecting mislabeled samples in vision models and selecting data samples for improving performance of natural language processing transformer models. We also extend its use to influential sample identification for fine-tuning Large Language Models. Anshuman Chhabra, Jian Chen 0016, Prasant Mohapatra, Hongfu Liu 0001 |
ICML | 5 |
| 2024 | "What Data Benefits My Classifier?" Enhancing Model Performance and Interpretability through Influence-Based Data SelectionabstractClassification models are ubiquitously deployed in society and necessitate high utility, fairness, and robustness performance. Current research efforts mainly focus on improving model architectures and learning algorithms on fixed datasets to achieve this goal. In contrast, in this paper, we address an orthogonal yet crucial problem: given a fixed convex learning model (or a convex surrogate for a non-convex model) and a function of interest, we assess what data benefits the model by interpreting the feature space, and then aim to improve performance as measured by this function. To this end, we propose the use of influence estimation models for interpreting the classifier's performance from the perspective of the data feature space. Additionally, we propose data selection approaches based on influence that enhance model utility, fairness, and robustness. Through extensive experiments on synthetic and real-world datasets, we validate and demonstrate the effectiveness of our approaches not only for conventional classification scenarios, but also under more challenging scenarios such as distribution shifts, fairness poisoning attacks, utility evasion attacks, online learning, and active learning. Anshuman Chhabra, Peizhao Li, Prasant Mohapatra, Hongfu Liu 0001 |
ICLR | 4 |
| 2024 | Category-Aware Active Domain AdaptationabstractActive domain adaptation has shown promising results in enhancing unsupervised domain adaptation (DA), by actively selecting and annotating a small amount of unlabeled samples from the target domain. Despite its effectiveness in boosting overall performance, the gain usually concentrates on the categories that are readily improvable, while challenging categories that demand the utmost attention are often overlooked by existing models. To alleviate this discrepancy, we propose a novel category-aware active DA method that aims to boost the adaptation for the individual category without adversely affecting others. Specifically, our approach identifies the unlabeled data that are most important for the recognition of the targeted category. Our method assesses the impact of each unlabeled sample on the recognition loss of the target data via the influence function, which allows us to directly evaluate the sample importance, without relying on indirect measurements used by existing methods. Comprehensive experiments and in-depth explorations demonstrate the efficacy of our method on category-aware active DA over three datasets. Wenxiao Xiao, Jiuxiang Gu, Hongfu Liu 0001 |
ICML | 3 |
| 2024 | Graph-Graph Similarity NetworkabstractGraph learning aims to predict the label for an entire graph. Recently, graph neural network (GNN)-based approaches become an essential strand to learning low-dimensional continuous embeddings of entire graphs for graph label prediction. While GNNs explicitly aggregate the neighborhood information and implicitly capture the topological structure for graph representation, they ignore the relationships among graphs. In this article, we propose a graph-graph (G2G) similarity network to tackle the graph learning problem by constructing a SuperGraph through learning the relationships among graphs. Each node in the SuperGraph represents an input graph, and the weights of edges denote the similarity between graphs. By this means, the graph learning task is then transformed into a classical node label propagation problem. Specifically, we use an adversarial autoencoder to align embeddings of all the graphs to a prior data distribution. After the alignment, we design the G2G similarity network to learn the similarity between graphs, which functions as the adjacency matrix of the SuperGraph. By running node label propagation algorithms on the SuperGraph, we can predict the labels of graphs. Experiments on five widely used classification benchmarks and four public regression benchmarks under a fair setting demonstrate the effectiveness of our method. Pengyu Hong, Hongfu Liu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | Characterizing the Influence of Graph Elements
Zizhang Chen, Peizhao Li, Hongfu Liu 0001, Pengyu Hong |
ICLR | 3 |
| 2023 | Robust Fair Clustering: A Novel Fairness Attack and Defense Framework
Anshuman Chhabra, Peizhao Li, Prasant Mohapatra, Hongfu Liu 0001 |
ICLR | 4 |
| 2023 | Learning Antidote Data to Individual UnfairnessabstractFairness is essential for machine learning systems deployed in high-stake applications. Among all fairness notions, individual fairness, deriving from a consensus that `similar individuals should be treated similarly,' is a vital notion to describe fair treatment for individual cases. Previous studies typically characterize individual fairness as a prediction-invariant problem when perturbing sensitive attributes on samples, and solve it by Distributionally Robust Optimization (DRO) paradigm. However, such adversarial perturbations along a direction covering sensitive information used in DRO do not consider the inherent feature correlations or innate data constraints, therefore could mislead the model to optimize at off-manifold and unrealistic samples. In light of this drawback, in this paper, we propose to learn and generate antidote data that approximately follows the data distribution to remedy individual unfairness. These generated on-manifold antidote data can be used through a generic optimization procedure along with original training data, resulting in a pure pre-processing approach to individual unfairness, or can also fit well with the in-processing DRO paradigm. Through extensive experiments on multiple tabular datasets, we demonstrate our method resists individual unfairness at a minimal or zero cost to predictive utility compared to baselines. Peizhao Li, Ethan Xia, Hongfu Liu 0001 |
ICML | 3 |
| 2023 | Transforming Complex Problems Into K-Means SolutionsabstractK-means is a fundamental clustering algorithm widely used in both academic and industrial applications. Its popularity can be attributed to its simplicity and efficiency. Studies show the equivalence of K-means to principal component analysis, non-negative matrix factorization, and spectral clustering. However, these studies focus on standard K-means with squared euclidean distance. In this review paper, we unify the available approaches in generalizing K-means to solve challenging and complex problems. We show that these generalizations can be seen from four aspects: data representation, distance measure, label assignment, and centroid updating. As concrete applications of transforming problems into modified K-means formulation, we review the following applications: iterative subspace projection and clustering, consensus clustering, constrained clustering, domain adaptation, and outlier detection. Hongfu Liu 0001, Junxiang Chen, Jennifer G. Dy, Yun Fu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | Second-Order Unsupervised Feature Selection via Knowledge Contrastive DistillationabstractUnsupervised feature selection aims to select a subset from the original features that are most useful for the downstream tasks without external guidance information. While most unsupervised feature selection methods focus on ranking features based on the intrinsic properties of data, most of them do not pay much attention to the relationships between features, which often leads to redundancy among the selected features. In this paper, we propose a two-stageSecond-Order unsupervisedFeature selection via knowledge contrastive disTillation (SOFT) model that incorporates the second-order covariance matrix with the first-order data matrix for unsupervised feature selection. In the first stage, we learn a sparse attention matrix that can represent second-order relations between features by contrastively distilling the intrinsic structure. In the second stage, we build a relational graph based on the learned attention matrix and perform graph segmentation. To this end, we conduct feature selection by only selecting one feature from each cluster to decrease the feature redundancy. Experimental results on 12 public datasets show that SOFT outperforms classical and recent state-of-the-art methods, which demonstrates the effectiveness of our proposed method. Moreover, we also provide rich in-depth experiments to further explore several key factors of SOFT. Jundong Li, Hongfu Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Algorithm 1038: KCC: A MATLAB Package for k-Means-based Consensus ClusteringabstractConsensus clustering is gaining increasing attention for its high quality and robustness. In particular, k -means-based Consensus Clustering (KCC) converts the usual computationally expensive problem to a classic k -means clustering with generalized utility functions, bringing potentials for large-scale data clustering on different types of data. Despite KCC’s applicability and generalizability, implementing this method such as representing the binary dataset in the k -means heuristic is challenging and has seldom been discussed in prior work. To fill this gap, we present a MATLAB package, KCC, that completely implements the KCC framework and utilizes a sparse representation technique to achieve a low space complexity. Compared to alternative consensus clustering packages, the KCC package is of high flexibility, efficiency, and effectiveness. Extensive numerical experiments are also included to show its usability on real-world datasets. Hao Lin 0002, Hongfu Liu 0001, Junjie Wu 0002, Stephan Günnemann |
ACM Trans. Math. Softw. | 2 |
| 2022 | Fairness of Machine Learning in Search EnginesabstractFairness has gained increasing importance in a variety of AI and machine learning contexts. As one of the most ubiquitous applications of machine learning, search engines mediate much of the information experiences of members of society. Consequently, understanding and mitigating potential algorithmic unfairness in search have become crucial for both users and systems. In this tutorial, we will introduce the fundamentals of fairness in machine learning, for both supervised learning such as classification and ranking, and unsupervised learning such as clustering. We will then present the existing work on fairness in search engines, including the fairness definitions, evaluation metrics, and taxonomies of methodologies. This tutorial will help orient information retrieval researchers to algorithmic fairness, provide an introduction to the growing literature on this topic, and gathering researchers and practitioners interested in this research direction. Yi Fang 0008, Hongfu Liu 0001, Zhiqiang Tao, Mikhail Yurochkin |
CIKM | 2 |
| 2022 | Exploiting Temporal Relations on Radar Perception for Autonomous DrivingabstractWe consider the object recognition problem in autonomous driving using automotive radar sensors. Comparing to Lidar sensors, radar is cost-effective and robust in all- weather conditions for perception in autonomous driving. However, radar signals suffer from low angular resolution and precision in recognizing surrounding objects. To enhance the capacity of automotive radar, in this work, we exploit the temporal information from successive ego-centric bird-eye-view radar image frames for radar object recognition. We leverage the consistency of an object's existence and attributes (size, orientation, etc.), and propose a temporal relational layer to explicitly model the relations between objects within successive radar images. In both object detection and multiple object tracking, we show the superiority of our method compared to several baseline approaches. Peizhao Li, Pu Wang 0004, Karl Berntorp, Hongfu Liu 0001 |
CVPR | 4 |
| 2022 | Achieving Fairness at No Utility Cost via Data Reweighing with InfluenceabstractWith the fast development of algorithmic governance, fairness has become a compulsory property for machine learning models to suppress unintentional discrimination. In this paper, we focus on the pre-processing aspect for achieving fairness, and propose a data reweighing approach that only adjusts the weight for samples in the training phase. Different from most previous reweighing methods which usually assign a uniform weight for each (sub)group, we granularly model the influence of each training sample with regard to fairness-related quantity and predictive utility, and compute individual weights based on influence under the constraints from both fairness and utility. Experimental results reveal that previous methods achieve fairness at a non-negligible cost of utility, while as a significant advantage, our approach can empirically release the tradeoff and obtain cost-free fairness for equal opportunity. We demonstrate the cost-free fairness through vanilla classifiers and standard training processes, compared to baseline methods on multiple real-world tabular datasets. Code available at https://github.com/brandeis-machine-learning/influence-fairness. Peizhao Li, Hongfu Liu 0001 |
ICML | 2 |
| 2022 | Multi-task Envisioning Transformer-based Autoencoder for Corporate Credit Rating Migration Early PredictionabstractCorporate credit ratings issued by third-party rating agencies are quantified assessments of a company's creditworthiness. Credit Ratings highly correlate to the likelihood of a company defaulting on its debt obligations. These ratings play critical roles in investment decision-making as one of the key risk factors. They are also central to the regulatory framework such as BASEL II in calculating necessary capital for financial institutions. Being able to predict rating changes will greatly benefit both investors and regulators alike. In this paper, we consider the corporate credit rating migration early prediction problem, which predicts the credit rating of an issuer will be upgraded, unchanged, or downgraded after 12 months based on its latest financial reporting information at the time. We investigate the effectiveness of different standard machine learning algorithms and conclude these models deliver inferior performance. As part of our contribution, we propose a new Multi-task Envisioning Transformer-based Autoencoder (META) model to tackle this challenging problem. META consists of Positional Encoding, Transformer-based Autoencoder, and Multi-task Prediction to learn effective representations for both migration prediction and rating prediction. This enables META to better explore the historical data in the training stage for one-year later prediction. Experimental results show that META outperforms all baseline models. Steve Q. Xia, Hongfu Liu 0001 |
KDD | 3 |
| 2022 | Label-invariant Augmentation for Semi-Supervised Graph ClassificationabstractRecently, contrastiveness-based augmentation surges a new climax in the computer vision domain, where some operations, including rotation, crop, and flip, combined with dedicated algorithms, dramatically increase the model generalization and robustness. Following this trend, some pioneering attempts employ the similar idea to graph data. Nevertheless, unlike images, it is much more difficult to design reasonable augmentations without changing the nature of graphs. Although exciting, the current graph contrastive learning does not achieve as promising performance as visual contrastive learning. We conjecture the current performance of graph contrastive learning might be limited by the violation of the label-invariant augmentation assumption. In light of this, we propose a label-invariant augmentation for graph-structured data to address this challenge. Different from the node/edge modification and subgraph extraction, we conduct the augmentation in the representation space and generate the augmented samples in the most difficult direction while keeping the label of augmented data the same as the original samples. In the semi-supervised scenario, we demonstrate our proposed method outperforms the classical graph neural network based methods and recent graph contrastive learning on eight benchmark graph-structured data, followed by several in-depth experiments to further explore the label-invariant augmentation in several aspects. Chuxu Zhang, Hongfu Liu 0001 |
NeurIPS | 4 |
| 2022 | A Gospel for MOBA Game: Ranking-Preserved Hero Change Prediction in Dota 2abstractDota 2is one of the most popular multiplayer online battle arena games, in which players of different teams controlling heroes fight each other to pursue a championship. To enrich the diversity of heroes and provide a balanced battle environment,Dota 2game designers keep changing the attributes or skills of heroes in constantly updated game versions. SinceDota 2is an intricate game, and numerous factors are involved in a match, it is challenging to figure out how to adjust heroes to meet the balance. As far as we know, no effort from intelligent learning perspective has been made to judge whether a hero should be changed in the next game version. This article proposes a ranking-preserved method to predict whether a hero should be enhanced, weakened, or unchanged. Specifically, a unified ranking-preserved temporal graph convolutional network model containing ranking preservation, graph convolutional network, and long short-term memory is designed to generate a ranking list of all the heroes, which indicates the strength of heroes as well as which heroes to be enhanced or weakened. The experiments on match records show that our approach provides a high-quality prediction and performs better than other baseline models. For game players, our model can give them a view of which heroes are more powerful currently and help them make better choices in choosing heroes. For game designers, our model can provide statistical supports and model interpretation for hero adjustment. Hongfu Liu 0001, Jian Chen 0016 |
IEEE Trans. Games | 2 |
| 2022 | Self-Guided Deep Multiview Subspace Clustering via Consensus Affinity RegularizationabstractMultiview subspace clustering (MVSC) leverages the complementary information among different views of multiview data and seeks a consensus subspace clustering result better than that using any individual view. Though proved effective in some cases, existing MVSC methods often obtain unsatisfactory results since they perform subspace analysis with raw features that are often of high dimensions and contain noises. To remedy this, we propose a self-guided deep multiview subspace clustering (SDMSC) model that performs joint deep feature embedding and subspace analysis. SDMSC comprehensively explores multiview data and strives to obtain a consensus data affinity relationship agreed by features from not only all views but also all intermediate embedding spaces. With more constraints being cast, the desirable data affinity relationship is supposed to be more reliably recovered. Besides, to secure effective deep feature embedding without label supervision, we propose to use the data affinity relationship obtained with raw features as the supervision signals to self-guide the embedding process. With this strategy, the risk that our deep clustering model being trapped in bad local minima is reduced, bringing us satisfactory clustering results in a higher possibility. The experiments on seven widely used datasets show the proposed method significantly outperforms the state-of-the-art clustering methods. Our code is available at https://github.com/kailigo/dmvsc.git. Kai Li 0012, Hongfu Liu 0001, Yulun Zhang 0001, Yun Fu 0001 |
IEEE Trans. Cybern. | 2 |
| 2022 | Learnable Subspace ClusteringabstractThis article studies the large-scale subspace clustering (LS2C) problem with millions of data points. Many popular subspace clustering methods cannot directly handle the LS2C problem although they have been considered to be state-of-the-art methods for small-scale data points. A simple reason is that these methods often choose all data points as a large dictionary to build huge coding models, which results in high time and space complexity. In this article, we develop a learnable subspace clustering paradigm to efficiently solve the LS2C problem. The key concept is to learn a parametric function to partition the high-dimensional subspaces into their underlying low-dimensional subspaces instead of the computationally demanding classical coding models. Moreover, we propose a unified, robust, predictive coding machine (RPCM) to learn the parametric function, which can be solved by an alternating minimization algorithm. Besides, we provide a bounded contraction analysis of the parametric function. To the best of our knowledge, this article is the first work to efficiently cluster millions of data points among the subspace clustering methods. Experiments on million-scale data sets verify that our paradigm outperforms the related state-of-the-art methods in both efficiency and effectiveness. Jun Li 0027, Hongfu Liu 0001, Zhiqiang Tao, Handong Zhao, Yun Fu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2021 | Fairness-Aware Unsupervised Feature SelectionabstractFeature selection is a prevalent data preprocessing paradigm for various learning tasks. Due to the expensive cost of acquiring supervision information, unsupervised feature selection sparks great interests recently. However, existing unsupervised feature selection algorithms do not have fairness considerations and suffer from a high risk of amplifying discrimination by selecting features that are over associated with protected attributes such as gender, race, and ethnicity. In this paper, we make an initial investigation of the fairness-aware unsupervised feature selection problem and develop a principled framework, which leverages kernel alignment to find a subset of high-quality features that can best preserve the information in the original feature space while being minimally correlated with protected attributes. Specifically, different from the mainstream in-processing debiasing methods, our proposed framework can be regarded as a model-agnostic debiasing strategy that eliminates biases and discrimination before downstream learning algorithms are involved. Experimental results on real-world datasets demonstrate that our framework achieves a good trade-off between feature utility and promoting feature fairness. Xiaoying Xing, Hongfu Liu 0001, Chen Chen 0022, Jundong Li |
CIKM | 2 |
| 2021 | SelfDoc: Self-Supervised Document Representation LearningabstractWe propose SelfDoc, a task-agnostic pre-training framework for document image understanding. Because documents are multimodal and are intended for sequential reading, our framework exploits the positional, textual, and visual information of every semantically meaningful component in a document, and it models the contextualization between each block of content. Unlike existing document pre-training models, our model is coarse-grained instead of treating individual words as input, therefore avoiding an overly fine-grained with excessive contextualization. Beyond that, we introduce cross-modal learning in the model pre-training phase to fully leverage multimodal information from unlabeled documents. For downstream usage, we propose a novel modality-adaptive attention mechanism for multimodal feature fusion by adaptively emphasizing language and vision signals. Our framework benefits from self-supervised pre-training on documents without requiring annotations by a feature masking training strategy. It achieves superior performance on multiple downstream tasks with significantly fewer document images used in the pre-training stage compared to previous works. Peizhao Li, Jiuxiang Gu, Jason Kuen, Vlad I. Morariu, Handong Zhao, Rajiv Jain, Varun Manjunatha, Hongfu Liu 0001 |
CVPR | 8 |
| 2021 | Spatially Constrained GAN for Face and Fashion SynthesisabstractImage synthesis has raised tremendous attention in both academic and industrial areas, especially for conditional and target-oriented image synthesis, such as criminal portrait and fashion design. The current studies have achieved encouraging results along this direction, but they mostly focus on class labels where spatial contents are randomly generated from latent vectors. Some recent studies have explored spatial constraints for generative models guided by semantic segmentation, but most of them are designed for scene generation and lack random variation. Such methods are not suitable for face or fashion image synthesis, where different images may share the same semantics. Different from all the current methods, we decouple the image synthesis task into three independent dimensions and propose a novel Spatially Constrained Generative Adversarial Network (SCGAN) to model it. SCGAN uses a simple yet effective way to decouple spatial constraints and attribute conditions from latent vectors, and treat them as additional controllable signals via a segmentor and a specially designed generator. Other unregulated contents are left to be generated from latent vectors. Experimentally, we provide both qualitative and quantitative results on CelebA and DeepFashion datasets to demonstrate that the proposed SCGAN is very effective in synthesizing spatially controllable and attribute-specific images with high visual quality and large variations. Our code is provided at https://github.com/jackyjsy/SCGAN. Songyao Jiang, Hongfu Liu 0001, Yue Wu 0008, Yun Fu 0001 |
FG | 2 |
| 2021 | Towards Novel Target Discovery Through Open-Set Domain AdaptationabstractOpen-set domain adaptation (OSDA) considers that the target domain contains samples from novel categories unobserved in external source domain. Unfortunately, existing OSDA methods always ignore the demand for the information of unseen categories and simply recognize them as "un-known" set without further explanation. This motivates us to understand the unknown categories more specifically by exploring the underlying structures and recovering their interpretable semantic attributes. In this paper, we propose a novel framework to accurately identify the seen categories in target domain, and effectively recover the semantic attributes for unseen categories. Specifically, structure preserving partial alignment is developed to recognize the seen categories through domain-invariant feature learning. Attribute propagation over visual graph is designed to smoothly transit attributes from seen to unseen categories via visual-semantic mapping. Moreover, two new cross-domain benchmarks are constructed to evaluate the proposed framework in the novel and practical challenge. Experimental results on open-set recognition and semantic recovery demonstrate the superiority of the proposed method over other compared baselines. Taotao Jing, Hongfu Liu 0001, Zhengming Ding |
ICCV | 2 |
| 2021 | On Dyadic Fairness: Exploring and Mitigating Bias in Graph Connections
Peizhao Li, Yifei Wang 0002, Han Zhao 0002, Pengyu Hong, Hongfu Liu 0001 |
ICLR | 5 |
| 2021 | Deep Clustering based Fair Outlier DetectionabstractIn this paper, we focus on the fairness issues regarding unsupervised outlier detection. Traditional algorithms, without a specific design for algorithmic fairness, could implicitly encode and propagate statistical bias in data and raise societal concerns. To correct such unfairness and deliver a fair set of potential outlier candidates, we propose Deep Clustering based Fair Outlier Detection (DCFOD) that learns a good representation for utility maximization while enforcing the learnable representation to be subgroup-invariant on the sensitive attribute. Considering the coupled and reciprocal nature between clustering and outlier detection, we leverage deep clustering to discover the intrinsic cluster structure and out-of-structure instances. Meanwhile, an adversarial training erases the sensitive pattern for instances for fairness adaptation. Technically, we propose an instance-level weighted representation learning strategy to enhance the joint deep clustering and outlier detection, where the dynamic weight module re-emphasizes contributions of likely-inliers while mitigating the negative impact from outliers. Demonstrated by experiments on eight datasets comparing to 17 outlier detection algorithms, our DCFOD method consistently achieves superior performance on both the outlier detection validity and two types of fairness notions in outlier detection. Hanyu Song, Peizhao Li, Hongfu Liu 0001 |
KDD | 3 |
| 2021 | Implicit Semantic Response Alignment for Partial Domain AdaptationabstractPartial Domain Adaptation (PDA) addresses the unsupervised domain adaptation problem where the target label space is a subset of the source label space. Most state-of-art PDA methods tackle the inconsistent label space by assigning weights to classes or individual samples, in an attempt to discard the source data that belongs to the irrelevant classes. However, we believe samples from those extra categories would still contain valuable information to promote positive transfer. In this paper, we propose the Implicit Semantic Response Alignment to explore the intrinsic relationships among different categories by applying a weighted schema on the feature level. Specifically, we design a class2vec module to extract the implicit semantic topics from the visual features. With an attention layer, we calculate the semantic response according to each implicit semantic topic. Then semantic responses of source and target data are aligned to retain the relevant information contained in multiple categories by weighting the features, instead of samples. Experiments on several cross-domain benchmark datasets demonstrate the effectiveness of our method over the state-of-the-art PDA methods. Moreover, we elaborate in-depth analyses to further explore implicit semantic alignment. Wenxiao Xiao, Zhengming Ding, Hongfu Liu 0001 |
NeurIPS | 3 |
| 2021 | Clustering With Outlier RemovalabstractCluster analysis and outlier detection are two continuously rising topics in data mining area, which in fact connect to each other deeply. Cluster structure is vulnerable to outliers; inversely, outliers are the points belonging to none of any clusters. Unfortunately, most existing studies do not notice the coupled relationship between these two tasks and handle them separately. In this article, we consider the joint cluster analysis and outlier detection problem, and propose the Clustering with Outlier Removal (COR) algorithm. Specifically, the original space is transformed into a binary space via generating basic partitions. We employ Holoentropy to measure the compactness of each cluster without involving several outlier candidates. To provide a neat and efficient solution, an auxiliary binary matrix is introduced so that COR completely and efficiently solves the challenging problem via a unified K-means— with theoretical supports. Extensive experimental results on numerous data sets in various domains demonstrate the effectiveness and efficiency of COR significantly over state-of-the-art methods in terms of cluster validity and outlier detection. Some key factors including the basic partition number and generation strategy in COR with an application on abnormal flight trajectory detection are further analyzed for practical use. Hongfu Liu 0001, Jun Li 0027, Yue Wu 0008, Yun Fu 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2020 | Deep Fair Clustering for Visual LearningabstractFair clustering aims to hide sensitive attributes during data partition by balancing the distribution of protected subgroups in each cluster. Existing work attempts to address this problem by reducing it to a classical balanced clustering with a constraint on the proportion of protected subgroups of the input space. However, the input space may limit the clustering performance, and so far only low-dimensional datasets have been considered. In light of these limitations, in this paper, we propose Deep Fair Clustering (DFC) to learn fair and clustering-favorable representations for clustering simultaneously. Our approach could effectively filter out sensitive attributes from representations, and also lead to representations that are amenable for the following cluster analysis. Theoretically, we show that our fairness constraint in DFC will not incur much loss in terms of several clustering metrics. Empirically, we provide extensive experimental demonstrations on four visual datasets to corroborate the superior performance of the proposed approach over existing fair clustering and deep clustering methods on both cluster validity and fairness criterion. Peizhao Li, Han Zhao 0002, Hongfu Liu 0001 |
CVPR | 3 |
| 2020 | Evolutive preference analysis with online consumer ratings
Hongfu Liu 0001, Bin Zhu 0017 |
Inf. Sci. | 2 |
| 2020 | Softly Associative Transfer Learning for Cross-Domain ClassificationabstractThe main challenge of cross-domain text classification is to train a classifier in a source domain while applying it to a different target domain. Many transfer learning-based algorithms, for example, dual transfer learning, triplex transfer learning, etc., have been proposed for cross-domain classification, by detecting a shared low-dimensional feature representation for both source and target domains. These methods, however, often assume that the word clusters matrix or the clusters association matrix as knowledge transferring bridges are exactly the same across different domains, which is actually unrealistic in real-world applications and, therefore, could degrade classification performance. In light of this, in this paper, we propose a softly associative transfer learning algorithm for cross-domain text classification. Specifically, we integrate two non-negative matrix tri-factorizations into a joint optimization framework, with approximate constraints on both word clusters matrices and clusters association matrices so as to allow proper diversity in knowledge transfer, and with another approximate constraint on class labels in source domains in order to handle noisy labels. An iterative algorithm is then proposed to solve the above problem, with its convergence verified theoretically and empirically. Extensive experimental results on various text datasets demonstrate the effectiveness of our algorithm, even with the presence of abundant state-of-the-art competitors. Deqing Wang 0001, Chenwei Lu, Junjie Wu 0002, Hongfu Liu 0001, Fuzhen Zhuang, Hui Zhang 0028 |
IEEE Trans. Cybern. | 4 |
| 2020 | Marginalized Multiview Ensemble ClusteringabstractMultiview clustering (MVC), which aims to explore the underlying cluster structure shared by multiview data, has drawn more research efforts in recent years. To exploit the complementary information among multiple views, existing methods mainly learn a common latent subspace or develop a certain loss across different views, while ignoring the higher level information such as basic partitions (BPs) generated by the single-view clustering algorithm. In light of this, we propose a novel marginalized multiview ensemble clustering (M2VEC) method in this paper. Specifically, we solve MVC in an EC way, which generates BPs for each view individually and seeks for a consensus one. By this means, we naturally leverage the complementary information of multiview data upon the same partition space. In order to boost the robustness of our approach, the marginalized denoising process is adopted to mimic the data corruptions and noises, which provides robust partition-level representations for each view by training a single-layer autoencoder. A low-rank and sparse decomposition is seamlessly incorporated into the denoising process to explicitly capture the consistency information and meanwhile compensate the distinctness between heterogeneous features. Spectral consensus graph partitioning is also involved by our model to make M2VEC as a unified optimization framework. Moreover, a multilayer M2VEC is eventually delivered in a stacked fashion to encapsulate nonlinearity into partition-level representations for handling complex data. Experimental results on eight real-world data sets show the efficacy of our approach compared with several state-of-the-art multiview and EC methods. We also showcase our method performs well with partial multiview data. Zhiqiang Tao, Hongfu Liu 0001, Sheng Li 0001, Zhengming Ding, Yun Fu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2019 | Structured and Sparse Annotations for Image Emotion Distribution LearningabstractLabel distribution learning methods effectively address the label ambiguity problem and have achieved great success in image emotion analysis. However, these methods ignore structured and sparse information naturally contained in the annotations of emotions. For example, emotions can be grouped and ordered due to their polarities and degrees. Meanwhile, emotions have the character of intensity and are reflected in different levels of sparse annotations. Motivated by these observations, we present a convolutional neural network based framework called Structured and Sparse annotations for image emotion Distribution Learning (SSDL) to tackle two challenges. In order to utilize structured annotations, the Earth Mover’s Distance is employed to calculate the minimal cost required to transform one distribution to another for ordered emotions and emotion groups. Combined with Kullback-Leibler divergence, we design the loss to penalize the mispredictions according to the dissimilarities of same emotions and different emotions simultaneously. Moreover, in order to handle sparse annotations, sparse regularization based on emotional intensity is adopted. Through combined loss and sparse regularization, SSDL could effectively leverage structured and sparse annotations for predicting emotion distribution. Experiment results demonstrate that our proposed SSDL significantly outperforms the state-of-the-art methods. Haitao Xiong, Hongfu Liu 0001, Bineng Zhong 0001, Yun Fu 0001 |
AAAI | 2 |
| 2019 | Marginalized Latent Semantic Encoder for Zero-Shot LearningabstractZero-shot learning has been well explored to precisely identify new unobserved classes through a visual-semantic function obtained from the existing objects. However, there exist two challenging obstacles: one is that the human-annotated semantics are insufficient to fully describe the visual samples; the other is the domain shift across existing and new classes. In this paper, we attempt to exploit the intrinsic relationship in the semantic manifold when given semantics are not enough to describe the visual objects, and enhance the generalization ability of the visual-semantic function with marginalized strategy. Specifically, we design a Marginalized Latent Semantic Encoder (MLSE), which is learned on the augmented seen visual features and the latent semantic representation. Meanwhile, latent semantics are discovered under an adaptive graph reconstruction scheme based on the provided semantics. Consequently, our proposed algorithm could enrich visual characteristics from seen classes, and well generalize to unobserved classes. Experimental results on zero-shot benchmarks demonstrate that the proposed model delivers superior performance over the state-of-the-art zero-shot learning approaches. Zhengming Ding, Hongfu Liu 0001 |
CVPR | 2 |
| 2019 | Adversarial Graph Embedding for Ensemble ClusteringabstractEnsemble clustering generally integrates basic partitions into a consensus one through a graph partitioning method, which, however, has two limitations: 1) it neglects to reuse original features; 2) obtaining consensus partition with learnable graph representations is still under-explored. In this paper, we propose a novel Adversarial Graph Auto-Encoders (AGAE) model to incorporate ensemble clustering into a deep graph embedding process. Specifically, graph convolutional network is adopted as probabilistic encoder to jointly integrate the information from feature content and consensus graph, and a simple inner product layer is used as decoder to reconstruct graph with the encoded latent variables (i.e., embedding representations). Moreover, we develop an adversarial regularizer to guide the network training with an adaptive partition-dependent prior. Experiments on eight real-world datasets are presented to show the effectiveness of AGAE over several state-of-the-art deep embedding and ensemble clustering methods. Zhiqiang Tao, Hongfu Liu 0001, Jun Li 0027, Yun Fu 0001 |
IJCAI | 2 |
| 2019 | Multi-View Saliency-Guided Clustering for Image CosegmentationabstractImage cosegmentation aims at extracting the common objects from multiple images simultaneously. Existing methods mainly solve cosegmentation via the pre-defined graph, which lacks flexibility and robustness to handle various visual patterns. Besides, similar backgrounds also confuse the identifying of the common foreground. To address these issues, we propose a novel Multi-view Saliency-Guided Clustering algorithm (MvSGC) for the image cosegmentation task. In our model, the unsupervised saliency prior is used as partition-level side information to guide the foreground clustering process. To achieve robustness to noises and missing observations, similarities on instance-level and partition-level are both considered. Specifically, a unified clustering model with cosine similarity is proposed to capture the intrinsic structure of data and keep partition result consistent with the side information. Moreover, we leverage multi-view weight learning to integrate multiple feature representations to further improve the robustness of our approach. A K-means-like optimization algorithm is developed to proceed the constrained clustering in a highly efficient way with theoretical support. Experimental results on three benchmark datasets (i.e., the iCoseg, MSRC and Internet image dataset) and one RGB-D image dataset demonstrate the superiority of applying our clustering method for image cosegmentation. Zhiqiang Tao, Hongfu Liu 0001, Huazhu Fu, Yun Fu 0001 |
IEEE Trans. Image Process. | 2 |
| 2019 | Robust Spectral Ensemble Clustering via Rank MinimizationabstractEnsemble Clustering (EC) is an important topic for data cluster analysis. It targets to integrate multiple Basic Partitions (BPs) of a particular dataset into a consensus partition. Among previous works, one promising and effective way is to transform EC as a graph partitioning problem on the co-association matrix, which is a pair-wise similarity matrix summarized by all the BPs in essence. However, most existing EC methods directly utilize the co-association matrix, yet without considering various noises (e.g., the disagreement between different BPs and the outliers) that may exist in it. These noises can impair the cluster structure of a co-association matrix, and thus mislead the final graph partitioning process. To address this challenge, we propose a novel Robust Spectral Ensemble Clustering (RSEC) algorithm in this article. Specifically, we learn low-rank representation (LRR) for the co-association matrix to uncover its cluster structure and handle the noises, and meanwhile, we perform spectral clustering with the learned representation to seek for a consensus partition. These two steps are jointly proceeded within a unified optimization framework. In particular, during the optimizing process, we leverage consensus partition to iteratively enhance the block-diagonal structure of LRR, in order to assist the graph partitioning. To solve RSEC, we first formulate it by using nuclear norm as a convex proxy to the rank function. Then, motivated by the recent advances in non-convex rank minimization, we further develop a non-convex model for RSEC and provide it a solution by the majorization--minimization Augmented Lagrange Multiplier algorithm. Experiments on 18 real-world datasets demonstrate the effectiveness of our algorithm compared with state-of-the-art methods. Moreover, several impact factors on the clustering performance of our approach are also explored extensively. Zhiqiang Tao, Hongfu Liu 0001, Sheng Li 0001, Zhengming Ding, Yun Fu 0001 |
ACM Trans. Knowl. Discov. Data | 2 |
| 2019 | Structure-Preserved Unsupervised Domain AdaptationabstractDomain adaptation has been a primal approach to addressing the issues by lack of labels in many data mining tasks. Although considerable efforts have been devoted to domain adaptation with promising results, most existing work learns a classifier on a source domain and then predicts the labels for target data, where only the instances near the boundary determine the hyperplane and the whole structure information is ignored. Moreover, little work has been done regarding to multi-source domain adaptation. To that end, we develop a novel unsupervised domain adaptation framework, which ensures the whole structure of source domains is preserved to guide the target structure learning in a semi-supervised clustering fashion. To our knowledge, this is the first time when the domain adaptation problem is re-formulated as a semi-supervised clustering problem with target labels as missing values. Furthermore, by introducing an augmented matrix, a non-trivial solution is designed, which can be exactly mapped into a K-means-like optimization problem with modified distance function and update rule for centroids in an efficient way. Extensive experiments on several widely-used databases show the substantial improvements of our proposed approach over the state-of-the-art methods. Hongfu Liu 0001, Ming Shao, Zhengming Ding, Yun Fu 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2019 | Feature Selection with Unsupervised Consensus GuidanceabstractMost of the unsupervised feature selection methods employ pseudo labels generated by clustering to guide the feature selection; however, noisy and irrelevant features degrade the cluster structure, which is ineffective to supervise feature selection. In light of this, we propose the Consensus Guided Unsupervised Feature Selection (CGUFS) framework, which introduces consensus clustering to generate pseudo labels for feature selection. Generally speaking, multiple diverse basic partitions are generated from the data and the consensus clustering is employed to provide the high-quality and robust partition to guide the feature selection in a one-step framework. In addition, complex constraints such as non-negative are removed due to the crisp indicators of consensus clustering. Based on the CGUFS framework, two formulations are put forward by using the utility function and co-association matrix, respectively, and we propose the (weighted) K-means-like optimization solution for efficient solutions with theoretical supports. Moreover, we extend the CGUFS framework to handle multi-view data feature selection. Extensive experiments on several singleview and multi-view data mining data sets in different domains demonstrate that our methods outperform the most recent state-ofthe-art work in terms of effectiveness and efficiency. Some important impact factors and model parameters within CGUFS are thoroughly discussed for practical use. Hongfu Liu 0001, Ming Shao, Yun Fu 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2018 | Predictive Coding Machine for Compressed Sensing and Image DenoisingabstractSparse and low rank coding has widely received much attention in machine learning, multimedia and computer vision. Unfortunately, expensive inference restricts the power of coding models in real-world applications, e.g., compressed sensing and image deblurring. In order to avoid the expensive inference, we propose a predictive coding machine (PCM) which aims to train a deep neural network (DNN) encoder to approximate the codes. By this means, a test sample can be fast approximated by the well-trained DNN. However, DNN leads PCM to be a non-convex and non-smooth optimization problem, which is extremely hard to solve. To address this challenge, we extend accelerated proximal gradient for PCM by steering gradient descent of DNN. To the best of our knowledge, we are the first to propose a gradient descent algorithm guided by accelerated proximal gradient for solving the PCM problem. Besides, a sufficient condition is provided to ensure the convergence to a critical point. Moreover, when the coding models are convex in PCM, the convergence rate O(1/(m2√t)) can be held in which m is the iteration number of accelerated proximal gradient, and t is the epoch of training DNN. Numerical results verify the promising advantages of PCM in terms of effectiveness, efficiency and robustness. Jun Li 0027, Hongfu Liu 0001, Yun Fu 0001 |
AAAI | 2 |
| 2018 | Fast Clustering with Flexible Balance ConstraintsabstractBalanced clustering aims at partitioning a dataset with roughly even cluster sizes while exploiting the intrinsic structure of the data. Despite attracting increased attention recently in both the academia and the industry, most existing balanced clustering algorithms still have high run time complexities that prevent them from being applied to large datasets. To cope with this challenge, we propose a Fast Clustering with Flexible balance Constraints FCFC, a simple, fast and effective clustering algorithm that can deal with flexible balance constraints. In essence, FCFC employs K-means as the core clustering algorithm and the cluster size variances as the penalty for imbalance. The objective function consists of the combined classical K-means clustering cost as well as the imbalance penalty. By exploiting a new insight of the second term, FCFC is able to employ an efficient K-means-like optimization procedure that can scale to big datasets. Furthermore, we also extend our model for multiple balance constraints with theoretical supports. Extensive experimental results show that our method exceeds several state-of-the-art methods by large margins in terms of efficiency and clustering quality. Finally, a real-world application for Bing search is provided, where data are organized in multiple machines with data size and query frequency balancing objectives. In the simulated scenario, our solution achieves the same fidelity score while reduces cost by 75% compared to the baseline method. Hongfu Liu 0001, Ziming Huang, Mingqin Li, Yun Fu 0001 |
IEEE BigData | 1 |
| 2018 | Kinship Classification through Latent Adaptive SubspaceabstractWe tackle the challenging kinship classification problem. Different from kinship verification, which tells two persons have certain kinship relation or not, kinship classification aims to identify the family that a person belongs to. Beyond age and appearance gap across parents and children, the difficulties of kinship classification lie in that any data of the children to be classified are unavailable in advance to help training. To handle this challenge, an auxiliary database with complete parents and children modalities is employed to uncover the parent-children latent knowledge. Specifically, we propose a Latent Adaptive Subspace learning (LAS) to uncover the shared knowledge between two modalities so that the unseen test children are implicitly modeled as latent factors for kinship classification. Moreover, person-wise and family-wise constraints are designed to enhance the individual similarity and couple the parents and children within families for discriminative features. Comprehensive experiments on two large kinship datasets show that the proposed algorithm can effectively inherit knowledge from different databases and modalities and achieve the state-of-the-art performance. Yue Wu 0008, Zhengming Ding, Hongfu Liu 0001, Joseph P. Robinson, Yun Fu 0001 |
FG | 3 |
| 2018 | Infinite ensemble clustering
Hongfu Liu 0001, Ming Shao, Sheng Li 0001, Yun Fu 0001 |
Data Min. Knowl. Discov. | 1 |
| 2018 | Improving face representation learning with center invariant loss
Yue Wu 0008, Hongfu Liu 0001, Jun Li 0027, Yun Fu 0001 |
Image Vis. Comput. | 2 |
| 2018 | Partition Level Constrained ClusteringabstractConstrained clustering uses pre-given knowledge to improve the clustering performance. Here we use a new constraint called partition level side information and propose the Partition Level Constrained Clustering (PLCC) framework, where only a small proportion of the data is given labels to guide the procedure of clustering. Our goal is to find a partition which captures the intrinsic structure from the data itself, and also agrees with the partition level side information. Then we derive the algorithm of partition level side information based on K-means and give its corresponding solution. Further, we extend it to handle multiple side information and design the algorithm of partition level side information for spectral clustering. Extensive experiments demonstrate the effectiveness and efficiency of our method compared to pairwise constrained clustering and ensemble clustering methods, even in the inconsistent cluster number setting, which verifies the superiority of partition level side information to pairwise constraints. Besides, our method has high robustness to noisy side information, and we also validate the performance of our method with multiple side information. Finally, the image cosegmentation application based on saliency-guided side information demonstrates the effectiveness of PLCC as a flexible framework in different domains, even with the unsupervised side information. Hongfu Liu 0001, Zhiqiang Tao, Yun Fu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2018 | Visual Kinship Recognition of Families in the WildabstractWe present the largest database for visual kinship recognition, Families In the Wild (FIW), with over 13,000 family photos of 1,000 family trees with 4-to-38 members. It took only a small team to build FIW with efficient labeling tools and work-flow. To extend FIW, we further improved upon this process with a novel semi-automatic labeling scheme that used annotated faces and unlabeled text metadata to discover labels, which were then used, along with existing FIW data, for the proposed clustering algorithm that generated label proposals for all newly added data-both processes are shared and compared in depth, showing great savings in time and human input required. Essentially, the clustering algorithm proposed is semi-supervised and uses labeled data to produce more accurate clusters. We statistically compare FIW to related datasets, which unarguably shows enormous gains in overall size and amount of information encapsulated in the labels. We benchmark two tasks, kinship verification and family classification, at scales incomparably larger than ever before. Pre-trained CNN models fine-tuned on FIW outscores other conventional methods and achieved state-of-the art on the renowned KinWild datasets. We also measure human performance on kinship recognition and compare to a fine-tuned CNN. Joseph P. Robinson, Ming Shao, Yue Wu 0008, Hongfu Liu 0001, Timothy Gillis, Yun Fu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2018 | Consensus Regularized Multi-View Outlier DetectionabstractIdentifying different types of data outliers with abnormal behaviors in multi-view data setting is challenging due to the complicated data distributions across different views. Conventional approaches achieve this by learning a new latent feature representation with the pairwise constraint on different view data. In this paper, we argue that the existing methods are expensive in generalizing their models from two-view data to three-view (or more) data, in terms of the number of introduced variables and detection performance. To address this, we propose a novel multi-view outlier detection method with consensus regularization on the latent representations. Specifically, we explicitly characterize each kind of outliers by the intrinsic cluster assignment labels and sample-specific errors. Moreover, we make a thorough discussion about the proposed consensus-regularization and the pairwise-regularization. Correspondingly, an optimization solution based on augmented Lagrangian multiplier method is proposed and derived in details. In the experiments, we evaluate our method on five well-known machine learning data sets with different outlier settings. Further, to show its effectiveness in real-world computer vision scenario, we tailor our proposed model to saliency detection and face reconstruction applications. The extensive results of both standard multi-view outlier detection task and the extended computer vision tasks demonstrate the effectiveness of our proposed method. Handong Zhao, Hongfu Liu 0001, Zhengming Ding, Yun Fu 0001 |
IEEE Trans. Image Process. | 2 |
| 2018 | Consensus Guided Multi-View ClusteringabstractIn recent decades, tremendous emerging techniques thrive the artificial intelligence field due to the increasing collected data captured from multiple sensors. These multi-view data provide more rich information than traditional single-view data. Fusing heterogeneous information for certain tasks is a core part of multi-view learning, especially for multi-view clustering. Although numerous multi-view clustering algorithms have been proposed, most scholars focus on finding the common space of different views, but unfortunately ignore the benefits from partition level by ensemble clustering. For ensemble clustering, however, there is no interaction between individual partitions from each view and the final consensus one. To fill the gap, we propose a Consensus Guided Multi-View Clustering (CMVC) framework, which incorporates the generation of basic partitions from each view and fusion of consensus clustering in an interactive way, i.e., the consensus clustering guides the generation of basic partitions, and high quality basic partitions positively contribute to the consensus clustering as well. We design a non-trivial optimization solution to formulate CMVC into two iterative k -means clusterings by an approximate calculation. In addition, the generalization of CMVC provides a rich feasibility for different scenarios, and the extension of CMVC with incomplete multi-view clustering further validates the effectiveness for real-world applications. Extensive experiments demonstrate the advantages of CMVC over other widely used multi-view clustering methods in terms of cluster validity, and the robustness of CMVC to some important parameters and incomplete multi-view data. Hongfu Liu 0001, Yun Fu 0001 |
ACM Trans. Knowl. Discov. Data | 1 |
| 2017 | Simultaneous Clustering and EnsembleabstractEnsemble Clustering (EC) has gained a great deal of attention throughout the fields of data mining and machine learning, since it emerged as an effective and robust clustering framework. Typically, EC methods try to fuse multiple basic partitions (BPs) into a consensus one, of which each BP is obtained by performing traditional clustering method on the same dataset. One promising direction for ensemble clustering is to derive pairwise similarity from BPs, and then transform it as a graph partition problem. However, these graph based methods may suffer from an information loss when computing the similarity between data points, because they only utilize the categorical data provided by multiple BPs, yet neglect rich information from raw features. This problem can badly undermine the underlying cluster structure in the original feature space, and thus degrade the clustering performance. In light of this, we propose a novel Simultaneous Clustering and Ensemble (SCE) framework to alleviate such detrimental effect, which employs the similarity matrix from raw features to enhance the co-association matrix summarized by multiple BPs. Two neat closed-form solutions given by eigenvalue decomposition are provided for SCE. Experiments conducted on 16 real-world datasets demonstrate the effectiveness of the proposed SCE over the traditional clustering and state-of-the-art ensemble clustering methods. Moreover, several impact factors that may affect our method are also explored extensively. Zhiqiang Tao, Hongfu Liu 0001, Yun Fu 0001 |
AAAI | 2 |
| 2017 | Image Cosegmentation via Saliency-Guided Constrained Clustering with Cosine SimilarityabstractCosegmentation jointly segments the common objects from multiple images. In this paper, a novel clustering algorithm, called Saliency-Guided Constrained Clustering approach with Cosine similarity (SGC3), is proposed for the image cosegmentation task, where the common foregrounds are extracted via a one-step clustering process. In our method, the unsupervised saliency prior is utilized as a partition-level side information to guide the clustering process. To guarantee the robustness to noise and outlier in the given prior, the similarities of instance-level and partition-level are jointly computed for cosegmentation. Specifically, we employ cosine distance to calculate the feature similarity between data point and its cluster centroid, and introduce a cosine utility function to measure the similarity between clustering result and the side information. These two parts are both based on the cosine similarity, which is able to capture the intrinsic structure of data, especially for the non-spherical cluster structure. Finally, a K-means-like optimization is designed to solve our objective function in an efficient way. Experimental results on two widely-used datasets demonstrate our approach achieves competitive performance over the state-of-the-art cosegmentation methods. Zhiqiang Tao, Hongfu Liu 0001, Huazhu Fu, Yun Fu 0001 |
AAAI | 2 |
| 2017 | Multi-view graph learning with adaptive label propagationabstractGraphs play an essential role in many data mining paradigms, such as semi-supervised classification. Conventional graph learning methods mainly focus on constructing graphs from single-view data. Nowadays data can be collected from multiple views using various sensors. How to construct a robust and reliable graph from multi-view data is still an open problem. In this paper, we propose a multi-view graph learning (MVGL) approach with adaptive label propagation for semi-supervised classification. MVGL integrates latent factor extraction, graph sparsification, and label propagation into a unified framework. It seeks shared latent factors from multi-view data as view-independent data representations, and then constructs a sparse graph accordingly. Meanwhile, the label propagation is adaptively optimized during graph construction. An efficient optimization algorithm is designed to solve the model. Experimental results on two benchmark datasets show remarkable improvements over both single-view and multi-view learning baselines. Sheng Li 0001, Hongfu Liu 0001, Zhiqiang Tao, Yun Fu 0001 |
IEEE BigData | 2 |
| 2017 | Projective Low-rank Subspace Clustering via Learning Deep EncoderabstractLow-rank subspace clustering (LRSC) has been considered as the state-of-the-art method on small datasets. LRSC constructs a desired similarity graph by low-rank representation (LRR), and employs a spectral clustering to segment the data samples. However, effectively applying LRSC into clustering big data becomes a challenge because both LRR and spectral clustering suffer from high computational cost. To address this challenge, we create a projective low-rank subspace clustering (PLrSC) scheme for large scale clustering problem. First, a small dataset is randomly sampled from big dataset. Second, our proposed predictive low-rank decomposition (PLD) is applied to train a deep encoder by using the small dataset, and the deep encoder is used to fast compute the low-rank representations of all data samples. Third, fast spectral clustering is employed to segment the representations. As a non-trivial contribution, we theoretically prove the deep encoder can universally approximate to the exact (or bounded) recovery of the row space. Experiments verify that our scheme outperforms the related methods on large scale datasets in a small amount of time. We achieve the state-of-art clustering accuracy by 95.8% on MNIST using scattering convolution features. Jun Li 0027, Hongfu Liu 0001, Handong Zhao, Yun Fu 0001 |
IJCAI | 2 |
| 2017 | From Ensemble Clustering to Multi-View ClusteringabstractMulti-View Clustering (MVC) aims to find the cluster structure shared by multiple views of a particular dataset. Existing MVC methods mainly integrate the raw data from different views, while ignoring the high-level information. Thus, their performance may degrade due to the conflict between heterogeneous features and the noises existing in each individual view. To overcome this problem, we propose a novel Multi-View Ensemble Clustering (MVEC) framework to solve MVC in an Ensemble Clustering (EC) way, which generates Basic Partitions (BPs) for each view individually and seeks for a consensus partition among all the BPs. By this means, we naturally leverage the complementary information of multi-view data in the same partition space. Instead of directly fusing BPs, we employ the low-rank and sparse decomposition to explicitly consider the connection between different views and detect the noises in each view. Moreover, the spectral ensemble clustering task is also involved by our framework with a carefully designed constraint, making MVEC a unified optimization framework to achieve the final consensus partition. Experimental results on six real-world datasets show the efficacy of our approach compared with both MVC and EC methods. Zhiqiang Tao, Hongfu Liu 0001, Sheng Li 0001, Zhengming Ding, Yun Fu 0001 |
IJCAI | 2 |
| 2017 | Entropy-based consensus clustering for patient stratificationabstractMOTIVATION: Patient stratification or disease subtyping is crucial for precision medicine and personalized treatment of complex diseases. The increasing availability of high-throughput molecular data provides a great opportunity for patient stratification. Many clustering methods have been employed to tackle this problem in a purely data-driven manner. Yet, existing methods leveraging high-throughput molecular data often suffers from various limitations, e.g. noise, data heterogeneity, high dimensionality or poor interpretability. RESULTS: Here we introduced an Entropy-based Consensus Clustering (ECC) method that overcomes those limitations all together. Our ECC method employs an entropy-based utility function to fuse many basic partitions to a consensus one that agrees with the basic ones as much as possible. Maximizing the utility function in ECC has a much more meaningful interpretation than any other consensus clustering methods. Moreover, we exactly map the complex utility maximization problem to the classic K -means clustering problem, which can then be efficiently solved with linear time and space complexity. Our ECC method can also naturally integrate multiple molecular data types measured from the same set of subjects, and easily handle missing values without any imputation. We applied ECC to 110 synthetic and 48 real datasets, including 35 cancer gene expression benchmark datasets and 13 cancer types with four molecular data types from The Cancer Genome Atlas. We found that ECC shows superior performance against existing clustering methods. Our results clearly demonstrate the power of ECC in clinically relevant patient stratification. AVAILABILITY AND IMPLEMENTATION: The Matlab package is available at http://scholar.harvard.edu/yyl/ecc . CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Hongfu Liu 0001, Hongsheng Fang, Feixiong Cheng, Yun Fu 0001, Yang-Yu Liu |
Bioinform. | 1 |
| 2017 | Fuzzy Consensus Clustering With Applications on Big DataabstractConsensus clustering aims to find a single partition of data that agrees as much as possible with existing basic partitions. Given its robustness and generalizability, consensus clustering has emerged as a promising solution to find cluster structures inside heterogeneous big data rising from various application domains. In the area of fuzzy systems, however, research along this line is still in its initial stage with some unsystematic algorithmic studies. Finding a fuzzy consensus partition from multiple fuzzy basic partitions in an efficient, flexible, and robust way is still an exciting open problem calling for further investigation. In light of this, this paper provides a systematic study of fuzzy consensus clustering (FCC) from a utility perspective. Specifically, we first define the objective function of FCC clearly using the novel fuzzified contingency matrix. We then derive a family of FCC Utility functions termed as FCCU that can transform FCC to a weighted piecewise fuzzy $c$ -means clustering (piFCM) problem. This helps us to establish an algorithmic framework for FCC with flexible choice of utility functions, and speeds FCC significantly with a FCM-like iterative process of piFCM. To meet the big data challenge, we further parallelize FCC on the Spark platform with both vertical and horizontal segmentation schemes. Extensive experiments on various real-world datasets demonstrate the excellent performance of FCC, even with a majority of poor basic partitions. In particular, our method exhibits interesting potential for big data clustering in two real-life applications concerned with online event detection and overlapping community detection, respectively. Junjie Wu 0002, Zhiang Wu 0001, Jie Cao 0001, Hongfu Liu 0001, Yanchun Zhang |
IEEE Trans. Fuzzy Syst. | 4 |
| 2017 | Spectral Ensemble Clustering via Weighted K-Means: Theoretical and Practical EvidenceabstractAs a promising way for heterogeneous data analytics, consensus clustering has attracted increasing attention in recent decades. Among various excellent solutions, the co-association matrix based methods form a landmark, which redefines consensus clustering as a graph partition problem. Nevertheless, the relatively high time and space complexities preclude it from wide real-life applications. We, therefore, propose Spectral Ensemble Clustering (SEC) to leverage the advantages of co-association matrix in information integration but run more efficiently. We disclose the theoretical equivalence between SEC and weighted K-means clustering, which dramatically reduces the algorithmic complexity. We also derive the latent consensus function of SEC, which to our best knowledge is the first to bridge co-association matrix based methods to the methods with explicit global objective functions. Further, we prove in theory that SEC holds the robustness, generalizability, and convergence properties. We finally extend SEC to meet the challenge arising from incomplete basic partitions, based on which a row-segmentation scheme for big data clustering is proposed. Experiments on various real-world data sets in both ensemble and multi-view clustering scenarios demonstrate the superiority of SEC to some state-of-the-art methods. In particular, SEC seems to be a promising candidate for big data clustering. Hongfu Liu 0001, Junjie Wu 0002, Tongliang Liu, Dacheng Tao, Yun Fu 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2016 | Consensus Guided Unsupervised Feature SelectionabstractFeature selection has been widely recognized as one of the key problems in data mining and machine learning community, especially for high-dimensional data with redundant information, partial noises and outliers. Recently, unsupervised feature selection attracts substantial research attentions since data acquisition is rather cheap today but labeling work is still expensive and time consuming. This is specifically useful for effective feature selection of clustering tasks. Recent works using sparse projection with pre-learned pseudo labels achieve appealing results; however, they generate pseudo labels with all features so that noisy and ineffective features degrade the cluster structure and further harm the performance of feature selection; besides, these methods suffer from complex composition of multiple constraints and computational inefficiency, e.g., eigen-decomposition. Differently, in this work we introduce consensus clustering for pseudo labeling, which gets rid of expensive eigen-decomposition and provides better clustering accuracy with high robustness. In addition, complex constraints such as non-negative are removed due to the crisp indicators of consensus clustering. Specifically, we propose one efficient formulation for our unsupervised feature selection by using the utility function and provide theoretical analysis on optimization rules and model convergence. Extensive experiments on several popular data sets demonstrate that our methods are superior to the most recent state-of-the-art works in terms of NMI. Hongfu Liu 0001, Ming Shao, Yun Fu 0001 |
AAAI | 1 |
| 2016 | Outlier detection via sampling ensembleabstractOutlier detection is a key technique in data ming and machine learning fields. The deviating characters of outliers make huge detrimental effects on the learning tasks. A lot of algorithms are therefore proposed to handle outliers from different perspectives, such as distance, density, angle and so on. Among these approaches, the density-based methods achieve better performance, but also suffer from huge time complexity. Recently, in order to accelerate the speed and improve the performance, the subsampling ensemble method attracts much attention, which has a reasonable theoretical interpretation and high performance. However, existing work only gives the partial picture of outlier detection via row-sampling, the effective portfolio of bi-sampling is still void. In light of this, we propose the general outlier detection framework via bi-sampling, Bi-Sampling Outlier Detection (BSOD) and provide the effective portfolios of the row and column-sampling ratios in a theoretical way. In addition, the benefits of BSOD are fully illustrated in terms of ensemble diversity and divide-and-conquer. Further we employ LOF within BSOD as BI-LOF to conduct extensive experiments. In general, on 30 synthetic and 17 real-world data sets we thoroughly explore the characteristics of BI-LOF with different numbers of instances, features, nearest neighbors, validate the theoretical analysis of BSOD condition on synthetic data sets, and show obvious advantages over other state-of-the-art algorithms in terms of low and high dimensional real-world data sets. And finally we use BI-LOF to conduct image outlier detection and show high quality and stableness of BI-LOF. Hongfu Liu 0001, Yun Fu 0001 |
IEEE BigData | 1 |
| 2016 | Robust Spectral Ensemble ClusteringabstractEnsemble Clustering (EC) aims to integrate multiple Basic Partitions (BPs) of the same dataset into a consensus one. It could be transformed as a graph partition problem on the co-association matrix derived from BPs. However, existing EC methods usually directly use the co-association matrix, yet without considering various noises (e.g., the disagreement between different BPs or outliers) that may exist in it. These noises can impair the cluster structure of a co-association matrix and thus degrade the final clustering performance. In this paper, we propose a novel Robust Spectral Ensemble Clustering (RSEC) approach to address this challenge. First, RSEC learns a robust representation for the co-association matrix through low-rank constraint, which reveals the cluster structure of a co-association matrix and captures various noises in it. Second, RSEC finds the consensus partition by conducting spectral clustering. These two steps are iteratively performed in a unified optimization framework. Most importantly, during our optimization process, we utilize consensus partition to iteratively enhance the block-diagonal structure of the learned representation to further assist the clustering process. Experiments on numerous real-world datasets demonstrate the effectiveness of our method compared with the state-of-the-art. Moreover, several impact factors that may affect the clustering performance of our approach are also explored extensively. Zhiqiang Tao, Hongfu Liu 0001, Sheng Li 0001, Yun Fu 0001 |
CIKM | 2 |
| 2016 | Robust Multi-View Feature SelectionabstractHigh-throughput technologies have enabled us to rapidly accumulate a wealth of diverse data types. These multi-view data contain much more information to uncover the cluster structure than single-view data, which draws raising attention in data mining and machine learning areas. On one hand, many features are extracted to provide enough information for better representations, on the other hand, such abundant features might result in noisy, redundant and irrelevant information, which harms the performance of the learning algorithms. In this paper, we focus on a new topic, multi-view unsupervised feature selection, which aims to discover the discriminative features in each view for better explanation and representation. Although there are some exploratory studies along this direction, most of them employ the traditional feature selection by putting the features in different views together and fail to evaluate the performance in the multi-view setting. The features selected in this way are difficult to explain due to the meaning of different views, which disobeys the goal of feature selection as well. In light of this, we intend to give a correct understanding of multi-view feature selection. Different from the existing work, which either incorrectly concatenates the features from different views, or takes huge time complexity to learn the pseudo labels, we propose a novel algorithm, Robust Multi-view Feature Selection (RMFS), which applies robust multi-view K-means to obtain the robust and high quality pseudo labels for sparse feature selection in an efficient way. Nontrivially we give the solution by taking the derivatives and further provide a K-means-like optimization to update several variables in a unified framework with the convergence guarantee. We demonstrate extensive experiments on three real-world multi-view data sets, which illustrate the effectiveness and efficiency of RMFS in terms of both single-view and multi-view evaluations by a large margin. Hongfu Liu 0001, Haiyi Mao, Yun Fu 0001 |
ICDM | 1 |
| 2016 | Structure-Preserved Multi-source Domain AdaptationabstractDomain adaptation has achieved promising results in many areas, such as image classification and object recognition. Although a lot of algorithms have been proposed to solve the task with different domain distributions, it remains a challenge for multi-source unsupervised domain adaptation. In addition, most of the existing algorithms learn a classifier on the source domain and predict the labels for the target data, which indicates that only the knowledge derived from the hyperplane is transferred to the target domain and the structure information is ignored. In light of this, we propose a novel algorithm for multi-source unsupervised domain adaptation. Generally speaking, we aim to preserve the whole structure from source domains and transfer it to serve the task on the target domain. The source and target data are put together for clustering, which simultaneously explores the structures of the source and target domains. The structure-preserved information from source domain further guides the clustering process on the target domain. Extensive experiments on two widely used databases on object recognition and face identification show the substantial improvement of our proposed approach over several state-of-the-art methods. Especially, our algorithm can take use of multi-source domains and achieve robust and better performance compared with the single source domain adaptation methods. Hongfu Liu 0001, Ming Shao, Yun Fu 0001 |
ICDM | 1 |
| 2016 | Incomplete Multi-Modal Visual Data Grouping
Handong Zhao, Hongfu Liu 0001, Yun Fu 0001 |
IJCAI | 2 |
| 2016 | Infinite Ensemble for Image ClusteringabstractImage clustering has been a critical preprocessing step for vision tasks, e.g., visual concept discovery, content-based image retrieval. Conventional image clustering methods use handcraft visual descriptors as basic features via K-means, or build the graph within spectral clustering. Recently, representation learning with deep structure shows appealing performance in unsupervised feature pre-treatment. However, few studies have discussed how to deploy deep representation learning to image clustering problems, especially the unified framework which integrates both representation learning and ensemble clustering for efficient image clustering still remains void. In addition, even though it is widely recognized that with the increasing number of basic partitions, ensemble clustering gets better performance and lower variances, the best number of basic partitions for a given data set is a pending problem. In light of this, we propose the Infinite Ensemble Clustering (IEC), which incorporates the power of deep representation and ensemble clustering in a one-step framework to fuse infinite basic partitions. Generally speaking, a set of basic partitions is firstly generated from the image data, then by converting the basic partitions to the 1-of-K codings, we link the marginalized auto-encoder to the infinite ensemble clustering with i.i.d. basic partitions, which can be approached by the closed-form solutions, finally we follow the layer-wise training procedure and feed the concatenated deep features to K-means for final clustering. Extensive experiments on diverse vision data sets with different levels of visual descriptors demonstrate both the time efficiency and superior performance of IEC compared to the state-of-the-art ensemble clustering and deep clustering methods. Hongfu Liu 0001, Ming Shao, Sheng Li 0001, Yun Fu 0001 |
KDD | 1 |
| 2015 | Clustering with Partition Level Side InformationabstractConstrained clustering uses pre-given knowledge to improve the clustering performance. Among existing literature, researchers usually focus on Must-Link and Cannot-Link pairwise constraints. However, pairwise constraints not only disobey the way we make decisions, but also suffer from the vulnerability of noisy constraints and the order of constraints. In light of this, we use partition level side information instead of pairwise constraints to guide the process of clustering. Compared with pairwise constraints, partition level side information keeps the consistency within partial structure and avoids self-contradictory and the impact of constraints order. Generally speaking, only small part of the data instances are given labels by human workers, which are used to supervise the procedure of clustering. Inspired by the success of ensemble clustering, we aim to find a clustering solution which captures the intrinsic structure from the data itself, and agrees with the partition level side information as much as possible. Then we derive the objective function and equivalently transfer it into a K-mean-like optimization problem. Extensive experiments on several real-world datasets demonstrate the effectiveness and efficiency of our method compared to pairwise constrained clustering and consensus clustering, which verifies the superiority of partition level side information to pairwise constraints. Besides, our method has high robustness to noisy side information. Hongfu Liu 0001, Yun Fu 0001 |
ICDM | 1 |
| 2015 | Spectral Ensemble ClusteringabstractEnsemble clustering, also known as consensus clustering, is emerging as a promising solution for multi-source and/or heterogeneous data clustering. The co-association matrix based method, which redefines the ensemble clustering problem as a classical graph partition problem, is a landmark method in this area. Nevertheless, the relatively high time and space complexity preclude it from real-life large-scale data clustering. We therefore propose SEC, an efficient Spectral Ensemble Clustering method based on co-association matrix. We show that SEC has theoretical equivalence to weighted K-means clustering and results in vastly reduced algorithmic complexity. We then derive the latent consensus function of SEC, which to our best knowledge is among the first to bridge co-association matrix based method to the methods with explicit object functions. The robustness and generalizability of SEC are then investigated to prove the superiority of SEC in theory. We finally extend SEC to meet the challenge rising from incomplete basic partitions, based on which a scheme for big data clustering can be formed. Experimental results on various real-world data sets demonstrate that SEC is an effective and efficient competitor to some state-of-the-art ensemble clustering methods and is also suitable for big data clustering. Hongfu Liu 0001, Tongliang Liu, Junjie Wu 0002, Dacheng Tao, Yun Fu 0001 |
KDD | 1 |
| 2015 | DIAS: A Disassemble-Assemble Framework for Highly Sparse Text ClusteringabstractUpon extensive studies, text clustering remains a critical challenge in data mining community. Even by various techniques proposed to overcome some of these challenges, there still exist problems when dealing with weakly related or even noisy features. In response to this, we propose a DIssemble-ASsemble (DIAS) framework for text clustering. DIAS employs simple random feature sampling to disassemble high-dimensional text data and gains diverse structural knowledge. This also does good to avoiding the bulk of noisy features. Then the multi-view knowledge is assembled by weighted Information-theoretic Consensus Clustering (ICC) in order to gain a high-quality consensus partitioning. Extensive experiments on eight real-world text data sets demonstrate the advantages of DIAS over other widely used methods. In particular, DIAS shows strengths in learning from very weak basic partitionings. In addition, it is the natural suitability to distributed computing that makes DIAS become a promising candidate for big text clustering. Hongfu Liu 0001, Junjie Wu 0002, Dacheng Tao, Yun Fu 0001 |
SDM | 1 |
| 2015 | K-Means-Based Consensus Clustering: A Unified ViewabstractThe objective of consensus clustering is to find a single partitioning which agrees as much as possible with existing basic partitionings. Consensus clustering emerges as a promising solution to find cluster structures from heterogeneous data. As an efficient approach for consensus clustering, the K-means based method has garnered attention in the literature, however the existing research efforts are still preliminary and fragmented. To that end, in this paper, we provide a systematic study of K-means-based consensus clustering (KCC). Specifically, we first reveal a necessary and sufficient condition for utility functions which work for KCC. This helps to establish a unified framework for KCC on both complete and incomplete data sets. Also, we investigate some important factors, such as the quality and diversity of basic partitionings, which may affect the performances of KCC. Experimental results on various realworld data sets demonstrate that KCC is highly efficient and is comparable to the state-of-the-art methods in terms of clustering quality. In addition, KCC shows high robustness to incomplete basic partitionings with many missing values. Junjie Wu 0002, Hongfu Liu 0001, Hui Xiong 0001, Jie Cao 0001, Jian Chen 0016 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2013 | How Many Zombies Around You?abstractRecent years have witnessed the explosive growth of online social media. Weibo, a famous "Chinese Twitter", has attracted over half billion users in less than four years. Among them are zombie users or bogus users, who are seemingly active common users but actually marionettes manipulated by intelligent software for economic interests. To probe such users thus becomes critically important for a healthy Weibo, but the existing studies along this line are still in initial stage due to the serious lack of labeled zombies and the limited attributes for user profiling. In light of this, in this paper, we figure out a commercial way for training set labeling, and propose a two-stage cascading model called ProZombie for zombie user recognition. ProZombie decomposes the training/predicting process into fast and refined phases in cascade, which greatly improves the modeling efficiency without sacrificing the accuracy. Moreover, 35 attributes including 16 newly proposed ones are employed for a panoramic description of Weibo users. Experiments on real-world labeled Weibo users demonstrate the effectiveness and efficiency of ProZombie. More interestingly, two case studies based on ProZombie successfully unveil the zombies hidden around common users, and their impact to information propagation on Weibo. To our best knowledge, this study is among the first to quantify these interesting observations on Weibo. Hongfu Liu 0001, Hao Lin 0002, Junjie Wu 0002, Zhiang Wu 0001 |
ICDM | 1 |
| 2013 | A Theoretic Framework of K-Means-Based Consensus Clustering
Junjie Wu 0002, Hongfu Liu 0001, Hui Xiong 0001, Jie Cao 0001 |
IJCAI | 2 |
| 2013 | SEA: a system for event analysis on chinese tweetsabstractRecent years have witnessed the explosive growth of online social media. Weibo, a famous "Chinese Twitter", has attracted over 0.5 billion users in less than four years, with more than 1000 tweets generated in every second. These tweets are informative but very fragmented, and thus would be better archived from an event perspective, as done by Weibo itself in the "Micro-Topic" program. This effort, however, is yet far from satisfaction for not providing enough analytical power to events. In light of this, in this demo paper, we propose SEA, a System for Event Analysis on Chinese tweets. In general, SEA is an event-centric, multi-functional platform that conducts panoramic analysis on Weibo events from various aspects, including the semantic information of the events, the temporal and spatial trends, the public sentiments, the hidden sub-events, the key users in the event diffusion and their preferences, etc. These functions are enabled by the integration of various analytical models and by the noSQL techniques adopted purposefully for massive tweets management. Finally, a case study on the "Spring Festival" event demonstrates the effectiveness of SEA. To our best knowledge, SEA is the first third-party system that provides panoramic analysis to Weibo events. Yaqiong Wang, Hongfu Liu 0001, Hao Lin 0002, Junjie Wu 0002, Zhiang Wu 0001, Jie Cao 0001 |
KDD | 2 |
| 2012 | Cosine interesting pattern discovery
Junjie Wu 0002, Shiwei Zhu, Hongfu Liu 0001, Guoping Xia |
Inf. Sci. | 3 |