EDBT 2026 Demo / reviewers in the wild / expert
Chun-Guang Li
dblp:75/1999
· DBLP profile ↗
52ranked-venue papers
8as first author
19since 2021 · last 2025
0000-0002-5716-268XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 36 · 7 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 30 · 6 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Temporal Rate Reduction Clustering for Human Motion SegmentationabstractHuman Motion Segmentation (HMS), which aims to partition videos into non-overlapping human motions, has attracted increasing research attention recently. Existing approaches for HMS are mainly dominated by subspace clustering methods, which are grounded on the assumption that high-dimensional temporal data align with a Union-of-Subspaces (UoS) distribution. However, the frames in video capturing complex human motions with cluttered backgrounds may not align well with the UoS distribution. In this paper, we propose a novel approach for HMS, named Temporal Rate Reduction Clustering ($\text{TR}^2\text{C}$), which jointly learns structured representations and affinity to segment the sequences of frames in video. Specifically, the structured representations learned by $\text{TR}^2\text{C}$ enjoy temporally consistency and are aligned well with a UoS structure, which is favorable for addressing the HMS task. We conduct extensive experiments on five benchmark HMS datasets and achieve state-of-the-art performances with different feature extractors. The code is available at: https://github.com/mengxianghan123/TR2C. Xianghan Meng, Zhengyu Tong, Chun-Guang Li |
ICCV | 4 |
| 2025 | Exploring a Principled Framework for Deep Subspace ClusteringabstractSubspace clustering is a classical unsupervised learning task, built on a basic assumption that high-dimensional data can be approximated by a union of subspaces (UoS). Nevertheless, the real-world data are often deviating from the UoS assumption. To address this challenge, state-of-the-art deep subspace clustering algorithms attempt to jointly learn UoS representations and self-expressive coefficients. However, the general framework of the existing algorithms suffers from feature collapse and lacks a theoretical guarantee to learn desired UoS representation. In this paper, we present a Principled fRamewOrk for Deep Subspace Clustering (PRO-DSC), which is designed to learn structured representations and self-expressive coefficients in a unified manner. Specifically, in PRO-DSC, we incorporate an effective regularization on the learned representations into the self-expressive model, prove that the regularized self-expressive model is able to prevent feature space collapse, and demonstrate that the learned optimal representations under certain condition lie on a union of orthogonal subspaces. Moreover, we provide a scalable and efficient approach to implement our PRO-DSC and conduct extensive experiments to verify our theoretical findings and demonstrate the superior performance of our proposed deep subspace clustering approach. Xianghan Meng, Xianbiao Qi, Rong Xiao 0003, Chun-Guang Li |
ICLR | 6 |
| 2025 | Taming Transformer Without Using Learning Rate WarmupabstractScaling Transformer to a large scale without using some technical tricks such as learning rate warump and an obviously lower learning rate, is an extremely challenging task, and is increasingly gaining more attention. In this paper, we provide a theoretical analysis for the process of training Transformer and reveal a key problem behind model crash phenomenon in the training process, termed *spectral energy concentration* of ${W_q}^{\top} W_k$, which is the reason for a malignant entropy collapse, where ${W_q}$ and $W_k$ are the projection matrices for the query and the key in Transformer, respectively.
To remedy this problem, motivated by *Weyl's Inequality*, we present a novel optimization strategy, \ie, making the weight updating in successive steps steady---if the ratio $\frac{\sigma_{1}(\nabla W_t)}{\sigma_{1}(W_{t-1})}$ is larger than a threshold, we will automatically bound the learning rate to a weighted multiple of $\frac{\sigma_{1}(W_{t-1})}{\sigma_{1}(\nabla W_t)}$, where $\nabla W_t$ is the updating quantity in step $t$. Such an optimization strategy can prevent spectral energy concentration to only a few directions, and thus can avoid malignant entropy collapse which will trigger the model crash. We conduct extensive experiments using ViT, Swin-Transformer and GPT, showing that our optimization strategy can effectively and stably train these (Transformer) models without using learning rate warmup. Xianbiao Qi, Yelin He, Jiaquan Ye, Chun-Guang Li, Bojia Zi, Xili Dai, Qin Zou 0001, Rong Xiao 0003 |
ICLR | 4 |
| 2025 | Towards Interpretable and Efficient Attention: Compressing All by Contracting a FewabstractAttention mechanisms have achieved significant empirical success in multiple fields, but their underlying optimization objectives remain unclear yet. Moreover, the quadratic complexity of self-attention has become increasingly prohibitive. Although interpretability and efficiency are two mutually reinforcing pursuits, prior work typically investigates them separately. In this paper, we propose a unified optimization objective that derives inherently interpretable and efficient attention mechanisms through algorithm unrolling. Precisely, we construct a gradient step of the proposed objective with a set of forward-pass operations of our \emph{Contract-and-Broadcast Self-Attention} (CBSA), which compresses input tokens towards low-dimensional structures by contracting a few representatives of them. This novel mechanism can not only scale linearly by fixing the number of representatives, but also covers the instantiations of varied attention mechanisms when using different sets of representatives. We conduct extensive experiments to demonstrate comparable performance and superior advantages over black-box attention mechanisms on visual tasks. Our work sheds light on the integration of interpretability and efficiency, as well as the unified formula of attention mechanisms. Code is available at \href{https://github.com/QishuaiWen/CBSA}{this https URL}. Qishuai Wen, Chun-Guang Li |
NeurIPS | 3 |
| 2025 | Variational Vision Transformer with Anti-Over-Smoothing Strategy for Robust Face Recognition
Yuying Zhao, Jiani Hu, Chun-Guang Li |
PRCV (7) | 3 |
| 2025 | On neural architecture search and hyperparameter optimization: A max-flow based approach
Chao Xue 0003, Xiaoxing Wang, Yibing Zhan, Junchi Yan, Chun-Guang Li |
Neural Networks | 6 |
| 2025 | Neural Normalized Cut: A differential and generalizable approach for spectral clustering
Shangzhi Zhang, Chun-Guang Li, Xianbiao Qi, Rong Xiao 0003, Jun Guo 0002 |
Pattern Recognit. | 3 |
| 2025 | A perturbed match filtering approach for face image quality assessment
Yuying Zhao, Mei Wang 0001, Jiani Hu, Weihong Deng, Chun-Guang Li |
Pattern Recognit. | 5 |
| 2024 | Graph Cut-Guided Maximal Coding Rate Reduction for Learning Image Embedding and Clustering
Xianghan Meng, Xianbiao Qi, Rong Xiao 0003, Chun-Guang Li |
ACCV (10) | 6 |
| 2024 | Rethinking Decoders for Transformer-based Semantic Segmentation: A Compression PerspectiveabstractState-of-the-art methods for Transformer-based semantic segmentation typically adopt Transformer decoders that are used to extract additional embeddings from image embeddings via cross-attention, refine either or both types of embeddings via self-attention, and project image embeddings onto the additional embeddings via dot-product. Despite their remarkable success, these empirical designs still lack theoretical justifications or interpretations, thus hindering potentially principled improvements. In this paper, we argue that there are fundamental connections between semantic segmentation and compression, especially between the Transformer decoders and Principal Component Analysis (PCA). From such a perspective, we derive a white-box, fully attentional DEcoder for PrIncipled semantiC segemenTation (DEPICT), with the interpretations as follows: 1) the self-attention operator refines image embeddings to construct an ideal principal subspace that aligns with the supervision and retains most information; 2) the cross-attention operator seeks to find a low-rank approximation of the refined image embeddings, which is expected to be a set of orthonormal bases of the principal subspace and corresponds to the predefined classes; 3) the dot-product operation yields compact representation for image embeddings as segmentation masks. Experiments conducted on dataset ADE20K find that DEPICT consistently outperforms its black-box counterpart, Segmenter, and it is light weight and more robust. Qishuai Wen, Chun-Guang Li |
NeurIPS | 2 |
| 2024 | Neural Architecture Selection as a Nash Equilibrium With Batch EntanglementabstractModeling the architecture search process on a supernet and applying a differentiable method to find the importance of architecture are among the leading tools for differentiable neural architectures search (DARTS). One fundamental problem in DARTS is how to discretize or select a single-path architecture from the pretrained one-shot architecture. Previous approaches mainly exploit heuristic or progressive search methods for discretization and selection, which are not efficient and easily trapped by local optimizations. To address these issues, we formulate the task of finding a proper single-path architecture as an architecture game among the edges and operations with the strategies "keep" and "drop" and show that the optimal one-shot architecture is a Nash equilibrium of the architecture game. Then, we propose a novel and effective approach for discretizing and selecting a proper single-path architecture, which is based on extracting the single-path architecture that associates the maximal coefficient of the Nash equilibrium with the strategy "keep" in the architecture game. To further improve the efficiency, we employ a mechanism of entangled Gaussian representation of mini-batches, inspired by the classic Parrondo's paradox. If some mini-batch formed uncompetitive strategies, the entanglement of mini-batches would ensure the games be combined and, thus, turn into strong ones. We conduct extensive experiments on benchmark datasets and demonstrate that our approach is significantly faster than the state-of-the-art progressive discretizing methods while maintaining competitive performance with higher maximum accuracy. Qian Li 0037, Chao Xue 0003, Chun-Guang Li, Chao Ma 0004, Xiaokang Yang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2022 | A Max-Flow Based Approach for Neural Architecture Search
Chao Xue 0003, Xiaoxing Wang, Junchi Yan, Chun-Guang Li |
ECCV (20) | 4 |
| 2022 | Learning graph normalization for graph neural networks
Xianbiao Qi, Chun-Guang Li, Rong Xiao 0003 |
Neurocomputing | 4 |
| 2022 | The devil in the tail: Cluster consolidation plus cluster adaptive balancing loss for unsupervised person re-identification
Chaoqun Lin, Chun-Guang Li, Jun Guo 0002 |
Pattern Recognit. | 4 |
| 2022 | Automated search space and search strategy selection for AutoML
Chao Xue 0003, Mengting Hu 0002, Xueqi Huang, Chun-Guang Li |
Pattern Recognit. | 4 |
| 2022 | EMU: Effective Multi-Hot Encoding Net for Lightweight Scene Text Recognition With a Large Character SetabstractDeploying a lightweight deep model for scene text recognition task on mobile devices has great commercial value. However, the conventional softmax-based one-hot classification module becomes a cumbersome obstacle when handling multi-languages or languages with large character set (e.g., Chinese) due to the rapid expansion of model parameters with the number of classes. To this end, we propose an Effective Multi-hot encoding and classification modUle (EMU) for scene text recognition in the scenario of multi-languages or languages with large character set. Specifically, EMU generates a binary multi-hot label for each class with a real-valued sub-network in training stage and produces the prediction by calculating the inner product between the multi-hot code and the multi-hot label. Compared to the softmax-based one-hot classifier, EMU reduces the storage requirement and the time cost in inference stage significantly, retaining similar performance. Furthermore, we design a convolution feature basedLightweight TransFormerto learn the effective features for EMU and consequently develop a lightweight scene text recognition framework, termedLight-Former-EMU. We conduct extensive experiments on seven public English benchmarks and two real-world Chinese challenge benchmarks. Experimental results verify the effectiveness of the proposed EMU and demonstrate the promising performance of the proposed Light-Former-EMU. Bingcong Li, Xianbiao Qi, Chun-Guang Li, Rong Xiao 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | Cluster-Guided Asymmetric Contrastive Learning for Unsupervised Person Re-IdentificationabstractUnsupervised person re-identification (Re-ID) aims to match pedestrian images from different camera views in an unsupervised setting. Existing methods for unsupervised person Re-ID are usually built upon the pseudo labels from clustering. However, the result of clustering depends heavily on the quality of the learned features, which are overwhelmingly dominated by colors in images. In this paper, we attempt to suppress the negative dominating influence of colors to learn more effective features for unsupervised person Re-ID. Specifically, we propose a Cluster-guided Asymmetric Contrastive Learning (CACL) approach for unsupervised person Re-ID, in which clustering result is leveraged to guide the feature learning in a properly designed asymmetric contrastive learning framework. In CACL, both instance-level and cluster-level contrastive learning are employed to help the siamese network learn discriminant features with respect to the clustering result within and between different data augmentation views, respectively. In addition, we also present a cluster refinement method, and validate that the cluster refinement step helps CACL significantly. Extensive experiments conducted on three benchmark datasets demonstrate the superior performance of our proposal. Chun-Guang Li, Jun Guo 0002 |
IEEE Trans. Image Process. | 2 |
| 2021 | Learning a Self-Expressive Network for Subspace ClusteringabstractState-of-the-art subspace clustering methods are based on the self-expressive model, which represents each data point as a linear combination of other data points. However, such methods are designed for a finite sample dataset and lack the ability to generalize to out-of-sample data. Moreover, since the number of self-expressive coefficients grows quadratically with the number of data points, their ability to handle large-scale datasets is often limited. In this paper, we propose a novel framework for subspace clustering, termed Self-Expressive Network (SENet), which employs a properly designed neural network to learn a self-expressive representation of the data. We show that our SENet can not only learn the self-expressive coefficients with desired properties on the training data, but also handle out-of-sample data. Besides, we show that SENet can also be leveraged to perform subspace clustering on large-scale datasets. Extensive experiments conducted on synthetic data and real world benchmark data validate the effectiveness of the proposed method. In particular, SENet yields highly competitive performance on MNIST, Fashion MNIST and Extended MNIST and state-of-the-art performance on CIFAR-10. Shangzhi Zhang, Chong You, René Vidal, Chun-Guang Li |
CVPR | 4 |
| 2021 | Rainy Night Scene Understanding With Near Scene Semantic AdaptationabstractDeep networks have been used for semantic segmentation tasks on scenes of outdoor environments with increasing popularity. However, the majority of existing work centers on daytime scenes with favorable illumination and weather conditions, and relies on supervision with pixel-level annotations. This paper seeks to address the problem of semantic segmentation for rainy, night-time scenes without using pixel-level annotations. We introduce a near scene semantic approach that uses images of daytime scenes as a bridge for transferring knowledge from pre-trained segmentation models to rainy night images. Specifically, we first present near scene oriented Representation Adaptation (RA) to reduce the domain shift on the representation level. Next, we adapt the segmentation model from the daytime scenario, under varying weather conditions, to the rainy night scenario by using near scene oriented Segmentation Space Adaptation (SSA). Consequently, this further reduces the impact of the domain shift on the segmentation space level. For evaluation, we created a new dataset containing 7000 distinct daytime-night-time image pairs of near scenes obtained by a webcam, and 5266 daytime-rainy night image pairs collected by a car-mounted camera. In addition, we carefully annotated 226 rainy night images with classes defined in Cityscapes. The experimental results clearly demonstrate the advantage of the proposed algorithm. Shuai Di, Chun-Guang Li, Honggang Zhang 0002, Semir Elezovikj, Chiu C. Tan 0001, Haibin Ling |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2020 | Stochastic Sparse Subspace ClusteringabstractState-of-the-art subspace clustering methods are based on self-expressive model, which represents each data point as a linear combination of other data points. By enforcing such representation to be sparse, sparse subspace clustering is guaranteed to produce a subspace-preserving data affinity where two points are connected only if they are from the same subspace. On the other hand, however, data points from the same subspace may not be well-connected, leading to the issue of over-segmentation. We introduce dropout to address the issue of over-segmentation, which is based on randomly dropping out data points in self-expressive model. In particular, we show that dropout is equivalent to adding a squared $\ell_2$ norm regularization on the representation coefficients, therefore induces denser solutions. Then, we reformulate the optimization problem as a consensus problem over a set of small-scale subproblems. This leads to a scalable and flexible sparse subspace clustering approach, termed Stochastic Sparse Subspace Clustering, which can effectively handle large scale datasets. Extensive experiments on synthetic data and real world datasets validate the efficiency and effectiveness of our proposal. Chun-Guang Li, Chong You |
CVPR | 2 |
| 2020 | Self-paced Bottom-up Clustering Network with Side Information for Person Re-IdentificationabstractPerson re-identification (Re-ID) has attracted a lot of research attention in recent years. However, supervised methods demand enormous amount of manually annotated data. In this paper, we propose a Self-Paced bottom-up Clustering Network with Side Information (SPCNet-SI) for unsupervised person Re-ID, where the side information comes from the serial number of the camera associated with each image. Specifically, our proposed SPCNet-SI exploits the camera side information to guide the feature learning and uses soft label in bottom-up clustering process, in which the camera association information is used in the repelled loss and the soft label based cluster information is used to select the candidate cluster pairs to merge. Moreover, a self-paced dynamic mechanism is developed to regularize the merging process such that the clustering is implemented in an easy-to-hard way with a slow-to-fast merging process. Experiments on two benchmark datasets Market-1501 and DukeMTMC-ReID demonstrate promising performance. Chun-Guang Li, Ruo-Pei Guo, Jun Guo 0002 |
ICPR | 2 |
| 2020 | Learning Convolution Feature Aggregation via Edge Attention Convolution Network for Person Re-IdentificationabstractPerson Re-Identification (Re-ID) is a challenging task of matching pedestrian images collected from nonoverlapping multiple camera views due to huge variations from pose changes, occlusions, varying illumination and clutter background. Recently, graph convolution network or graph neural network increasingly gains a lot of research attention in person Re-ID. However, the existing methods have not fully exploit the available features on the graph. In this paper, we propose an efficient and effective end-to-end trainable framework, termed Edge Attention Convolution Network (EACN), to perform convolution feature learning and attentive feature aggregation for person Re-ID, in which the learned convolution features on vertex and its edges are attentively aggregated on a dynamic graph. We conduct extensive experiments on two large benchmark datasets, Market-1501 and DukeMTMC. Experimental results validate the efficiency and effectiveness of our proposal. Chaoqun Lin, Ruo-Pei Guo, Xianbiao Qi, Chun-Guang Li |
VCIP | 5 |
| 2020 | Density-adaptive kernel based efficient reranking approaches for person reidentification
Ruo-Pei Guo, Chun-Guang Li, Jiaru Lin, Jun Guo 0002 |
Neurocomputing | 2 |
| 2020 | A novel Pooling Block for improving lightweight deep neural networks
Bo Xiao 0006, Chun-Guang Li, Qianfang Xu |
Pattern Recognit. Lett. | 3 |
| 2019 | Self-Supervised Convolutional Subspace Clustering NetworkabstractSubspace clustering methods based on data self-expression have become very popular for learning from data that lie in a union of low-dimensional linear subspaces. However, the applicability of subspace clustering has been limited because practical visual data in raw form do not necessarily lie in such linear subspaces. On the other hand, while Convolutional Neural Network (ConvNet) has been demonstrated to be a powerful tool for extracting discriminative features from visual data, training such a ConvNet usually requires a large amount of labeled data, which are unavailable in subspace clustering applications. To achieve simultaneous feature learning and subspace clustering, we propose an end-to-end trainable framework, called Self-Supervised Convolutional Subspace Clustering Network (S$^2$ConvSCN), that combines a ConvNet module (for feature learning), a self-expression module (for subspace clustering) and a spectral clustering module (for self-supervision) into a joint optimization framework. Particularly, we introduce a dual self-supervision that exploits the output of spectral clustering to supervise the training of the feature learning module (via a classification loss) and the self-expression module (via a spectral clustering loss). Our experiments on four benchmark datasets show the effectiveness of the dual self-supervision and demonstrate superior performance of our proposed approach. Chun-Guang Li, Chong You, Xianbiao Qi, Honggang Zhang 0002, Jun Guo 0002, Zhouchen Lin |
CVPR | 2 |
| 2019 | Is an Affine Constraint Needed for Affine Subspace Clustering?abstractSubspace clustering methods based on expressing each data point as a linear combination of other data points have achieved great success in computer vision applications such as motion segmentation, face and digit clustering. In face clustering, the subspaces are linear and subspace clustering methods can be applied directly. In motion segmentation, the subspaces are affine and an additional affine constraint on the coefficients is often enforced. However, since affine subspaces can always be embedded into linear subspaces of one extra dimension, it is unclear if the affine constraint is really necessary. This paper shows, both theoretically and empirically, that when the dimension of the ambient space is high relative to the sum of the dimensions of the affine subspaces, the affine constraint has a negligible effect on clustering performance. Specifically, our analysis provides conditions that guarantee the correctness of affine subspace clustering methods both with and without the affine constraint, and shows that these conditions are satisfied for high-dimensional data. Underlying our analysis is the notion of affinely independent subspaces, which not only provides geometrically interpretable correctness conditions, but also clarifies the relationships between existing results for affine subspace clustering. Chong You, Chun-Guang Li, Daniel P. Robinson, René Vidal |
ICCV | 2 |
| 2019 | Local Convex Representation with Pruning for Manifold ClusteringabstractHigh-dimensional data in many applications can be considered as samples drawn from a union of multiple lowdimensional manifolds. Assigning data points into their own manifolds is referred to manifold clustering. Inspired by recent advances in subspace clustering, in this paper, we present an efficient approach for manifold clustering, called Local Convex Representation (LCR), in which each data point is represented as a convex combination of other points in the local neighborhood and under some mild conditions the nonzero coefficients are guaranteed to correspond to the data points lying on the same manifold. Moreover, we incorporate the estimated intrinsic dimension of the manifold to prune the minor nonzero coefficients and validate that the pruning step helps LCR yield remarkable improvements. Experiments on synthetic data as well as real world data demonstrate promising performance. Chun-Guang Li |
VCIP | 2 |
| 2018 | Dual Attention Matching Network for Context-Aware Feature Sequence Based Person Re-IdentificationabstractTypical person re-identification (ReID) methods usually describe each pedestrian with a single feature vector and match them in a task-specific metric space. However, the methods based on a single feature vector are not sufficient enough to overcome visual ambiguity, which frequently occurs in real scenario. In this paper, we propose a novel end-to-end trainable framework, called Dual ATtention Matching network (DuATM), to learn context-aware feature sequences and perform attentive sequence comparison simultaneously. The core component of our DuATM framework is a dual attention mechanism, in which both intrasequence and inter-sequence attention strategies are used for feature refinement and feature-pair alignment, respectively. Thus, detailed visual cues contained in the intermediate feature sequences can be automatically exploited and properly compared. We train the proposed DuATM network as a siamese network via a triplet loss assisted with a decorrelation loss and a cross-entropy loss. We conduct extensive experiments on both image and video based ReID benchmark datasets. Experimental results demonstrate the significant advantages of our approach compared to the state-of-the-art methods. Jianlou Si, Honggang Zhang 0002, Chun-Guang Li, Jason Kuen, Xiangfei Kong, Alex Chichung Kot, Gang Wang 0012 |
CVPR | 3 |
| 2018 | Density-Adaptive Kernel based Re-Ranking for Person Re-IdentificationabstractPerson Re-Identification (ReID) refers to the task of verifying the identity of a pedestrian observed from nonoverlapping surveillance cameras views. Recently, it has been validated that re-ranking could bring extra performance improvements in person ReID. However, the current re-ranking approaches either require feedbacks from users or suffer from burdensome computation cost. In this paper, we propose to exploit a density-adaptive kernel technique to perform efficient and effective re-ranking for person ReID. Specifically, we present two simple yet effective re-ranking methods, termed inverse Density-Adaptive Kernel based Re-ranking (inv-DAKR) and bidirectional Density-Adaptive Kernel based Re-ranking (bi-DAKR), which are based on a smooth kernel function with a density-adaptive parameter. Experiments on six benchmark data sets confirm that our proposals are effective and efficient. Ruo-Pei Guo, Chun-Guang Li, Jiaru Lin |
ICPR | 2 |
| 2018 | Constrained Sparse Subspace Clustering with Side-InformationabstractSubspace clustering refers to the problem of segmenting high dimensional data drawn from a union of subspaces into the respective subspaces. In some applications, partial side-information to indicate “must-link” or “cannot-link” in clustering is available. This leads to the task of subspace clustering with side-information. However, in prior work the supervision value of the side-information for subspace clustering has not been fully exploited. To this end, in this paper, we present an enhanced approach for constrained subspace clustering with side-information, termed Constrained Sparse Subspace Clustering plus (CSSC+), in which the side-information is used not only in the stage of learning an affinity matrix but also in the stage of spectral clustering. Moreover, we propose to estimate clustering accuracy based on the partial side-information and theoretically justify the connection to the ground-truth clustering accuracy in terms of the Rand index. We conduct experiments on three cancer gene expression datasets to validate the effectiveness of our proposals. Chun-Guang Li, Jun Guo 0002 |
ICPR | 1 |
| 2018 | Cross-Domain Traffic Scene Understanding: A Dense Correspondence-Based Transfer Learning ApproachabstractUnderstanding traffic scene images taken from vehicle mounted cameras is important for high-level tasks, such as advanced driver assistance systems and autonomous driving. It is a challenging problem due to large variations under different weather or illumination conditions. In this paper, we tackle the problem of traffic scene understanding from a cross-domain perspective. We attempt to understand the traffic scene from images taken from the same location but under different weather or illumination conditions (e.g., understanding the same traffic scene from images on a rainy night with the help of images taken on a sunny day). To this end, we propose a dense correspondence-based transfer learning (DCTL) approach, which consists of three main steps: 1) extracting deep representations of traffic scene images via a fine-tuned convolutional neural network; 2) constructing compact and effective representations via cross-domain metric learning and subspace alignment for cross-domain retrieval; and 3) transferring the annotations from the retrieved best matching image to the test image based on cross-domain dense correspondences and a probabilistic Markov random field. To verify the effectiveness of our DCTL approach, we conduct extensive experiments on a challenging data set, which contains 1828 images from six weather or illumination conditions. Shuai Di, Honggang Zhang 0002, Chun-Guang Li, Xue Mei, Danil V. Prokhorov, Haibin Ling |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2018 | Spatial Pyramid-Based Statistical Features for Person Re-Identification: A Comprehensive EvaluationabstractPerson re-identification (Re-Id) across nonoverlapping camera views is one of challenging problems in surveillance video analysis. The difficulties in person Re-Id mainly come from the large appearance variations caused by camera view angle, human pose, illumination, and occlusion. Recently, extensive efforts have been cast into addressing this problem by developing invariant features or discriminative distance metrics. However, there is still a lack of systematic evaluations on the pipeline for feature extraction and combination. In this paper, we propose a spatial pyramid-based statistical feature extraction framework as a unified pipeline of feature extraction and combination for person Re-Id, and systematically evaluate the configuration details in feature extraction and the fusion strategies in feature combination. Extensive experiments on benchmark datasets demonstrate the critical components in feature extraction. Moreover, by combining multiple features, our proposed approach can yield state-of-the-art performance. It should be mentioned that our approach achieves rank 1 matching rate of 45.8% on dataset VIPeR and 61.5% on dataset CUHK01, respectively. Jianlou Si, Honggang Zhang 0002, Chun-Guang Li, Jun Guo 0002 |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2017 | Structured Sparse Subspace Clustering: A Joint Affinity Learning and Subspace Clustering FrameworkabstractSubspace clustering refers to the problem of segmenting data drawn from a union of subspaces. State-of-the-art approaches for solving this problem follow a two-stage approach. In the first step, an affinity matrix is learned from the data using sparse or low-rank minimization techniques. In the second step, the segmentation is found by applying spectral clustering to this affinity. While this approach has led to the state-of-the-art results in many applications, it is suboptimal, because it does not exploit the fact that the affinity and the segmentation depend on each other. In this paper, we propose a joint optimization framework - Structured Sparse Subspace Clustering (S3C) - for learning both the affinity and the segmentation. The proposed S3C framework is based on expressing each data point as a structured sparse linear combination of all other data points, where the structure is induced by a norm that depends on the unknown segmentation. Moreover, we extend the proposed S3C framework into Constrained S3C (CS3C) in which available partial side-information is incorporated into the stage of learning the affinity. We show that both the structured sparse representation and the segmentation can be found via a combination of an alternating direction method of multipliers with spectral clustering. Experiments on a synthetic data set, the Extended Yale B face data set, the Hopkins 155 motion segmentation database, and three cancer data sets demonstrate the effectiveness of our approach. Chun-Guang Li, Chong You, René Vidal |
IEEE Trans. Image Process. | 1 |
| 2017 | HEp-2 Cell Classification via Combining Multiresolution Co-Occurrence Texture and Large Region Shape InformationabstractIndirect immunofluorescence imaging of human epithelial type 2 (HEp-2) cell image is an effective evidence to diagnose autoimmune diseases. Recently, computer-aided diagnosis of autoimmune diseases by the HEp-2 cell classification has attracted great attention. However, the HEp-2 cell classification task is quite challenging due to large intraclass and small interclass variations. In this paper, we propose an effective approach for the automatic HEp-2 cell classification by combining multiresolution co-occurrence texture and large regional shape information. To be more specific, we propose to: 1) capture multiresolution co-occurrence texture information by a novel pairwise rotation-invariant co-occurrence of local Gabor binary pattern descriptor; 2) depict large regional shape information by using an improved Fisher vector model with RootSIFT features, which are sampled from large image patches in multiple scales; and 3) combine both features. We evaluate systematically the proposed approach on the IEEE International Conference on Pattern Recognition (ICPR) 2012, the IEEE International Conference on Image Processing (ICIP) 2013, and the ICPR 2014 contest datasets. The proposed method based on the combination of the introduced two features outperforms the winners of the ICPR 2012 contest using the same experimental protocol. Our method also greatly improves the winner of the ICIP 2013 contest under four different experimental setups. Using the leave-one-specimen-out evaluation strategy, our method achieves comparable performance with the winner of the ICPR 2014 contest that combined four features. Xianbiao Qi, Guoying Zhao 0001, Chun-Guang Li, Jun Guo 0002, Matti Pietikäinen |
IEEE J. Biomed. Health Informatics | 3 |
| 2016 | Oracle Based Active Set Algorithm for Scalable Elastic Net Subspace ClusteringabstractState-of-the-art subspace clustering methods are based on expressing each data point as a linear combination of other data points while regularizing the matrix of coefficients with ℓ1, ℓ2or nuclear norms. ℓ1regularization is guaranteed to give a subspace-preserving affinity (i.e., there are no connections between points from different subspaces) under broad theoretical conditions, but the clusters may not be connected. ℓ2and nuclear norm regularization often improve connectivity, but give a subspace-preserving affinity only for independent subspaces. Mixed ℓ1, ℓ2and nuclear norm regularizations offer a balance between the subspace-preserving and connectedness properties, but this comes at the cost of increased computational complexity. This paper studies the geometry of the elastic net regularizer (a mixture of the ℓ1and ℓ2norms) and uses it to derive a provably correct and scalable active set method for finding the optimal coefficients. Our geometric analysis also provides a theoretical justification and a geometric interpretation for the balance between the connectedness (due to ℓ2regularization) and subspace-preserving (due to ℓ1regularization) properties for elastic net subspace clustering. Our experiments show that the proposed active set method not only achieves state-of-the-art clustering performance, but also efficiently handles large-scale datasets. Chong You, Chun-Guang Li, Daniel P. Robinson, René Vidal |
CVPR | 2 |
| 2016 | Low-rank and structured sparse subspace clusteringabstractHigh dimensional data often lie approximately in low dimensional subspaces corresponding to multiple classes or categories. Segmenting the high dimensional data into their corresponding low dimensional subspaces is referred as subspace clustering. State of the art methods solve this problem in two steps. First, an affinity matrix is built from data based on self-expressiveness model, in which each data point is expressed as a linear combination of other data points. Second, the segmentation is obtained by spectral clustering. However, solving two dependent steps separately is still suboptimal. In this paper, we propose a joint affinity learning and spectral clustering approach for low-rank representation based subspace clustering, termed Low-Rank and Structured Sparse Subspace Clustering (LRS3C), where a subspace structured norm that depends on subspace clustering result is introduced into the objective of low-rank representation problem. We solve it efficiently via a combination of Linearized Alternation Direction Method (LADM) with spectral clustering. Experiments on Hopkins 155 motion segmentation database and Extended Yale B data set demonstrated the effectiveness of our method. Chun-Guang Li, Honggang Zhang 0002, Jun Guo 0002 |
VCIP | 2 |
| 2016 | Dynamic texture and scene classification by transferring deep image features
Xianbiao Qi, Chun-Guang Li, Guoying Zhao 0001, Xiaopeng Hong, Matti Pietikäinen |
Neurocomputing | 2 |
| 2015 | Structured Sparse Subspace Clustering: A unified optimization frameworkabstractSubspace clustering refers to the problem of segmenting data drawn from a union of subspaces. State of the art approaches for solving this problem follow a two-stage approach. In the first step, an affinity matrix is learned from the data using sparse or low-rank minimization techniques. In the second step, the segmentation is found by applying spectral clustering to this affinity. While this approach has led to state of the art results in many applications, it is sub-optimal because it does not exploit the fact that the affinity and the segmentation depend on each other. In this paper, we propose a unified optimization framework for learning both the affinity and the segmentation. Our framework is based on expressing each data point as a structured sparse linear combination of all other data points, where the structure is induced by a norm that depends on the unknown segmentation. We show that both the segmentation and the structured sparse representation can be found via a combination of an alternating direction method of multipliers with spectral clustering. Experiments on a synthetic data set, the Hopkins 155 motion segmentation database, and the Extended Yale B data set demonstrate the effectiveness of our approach. Chun-Guang Li, René Vidal |
CVPR | 1 |
| 2015 | Learning Semi-Supervised Representation Towards a Unified Optimization Framework for Semi-Supervised LearningabstractState of the art approaches for Semi-Supervised Learning (SSL) usually follow a two-stage framework -- constructing an affinity matrix from the data and then propagating the partial labels on this affinity matrix to infer those unknown labels. While such a two-stage framework has been successful in many applications, solving two subproblems separately only once is still suboptimal because it does not fully exploit the correlation between the affinity and the labels. In this paper, we formulate the two stages of SSL into a unified optimization framework, which learns both the affinity matrix and the unknown labels simultaneously. In the unified framework, both the given labels and the estimated labels are used to learn the affinity matrix and to infer the unknown labels. We solve the unified optimization problem via an alternating direction method of multipliers combined with label propagation. Extensive experiments on a synthetic data set and several benchmark data sets demonstrate the effectiveness of our approach. Chun-Guang Li, Zhouchen Lin, Honggang Zhang 0002, Jun Guo 0002 |
ICCV | 1 |
| 2015 | Regularization in metric learning for person re-identificationabstractMetric learning plays a critical role in person re-identification problem. Unfortunately, due to the small size of training data, the metric learning used in this scenario suffers from over-fitting which leads to degenerated performance. In this paper, we investigate the effect of regularization in metric learning for person re-identification. Concretely we formulate the distance function from three perspectives and hence present four different regularized metric learning methods. Experiments on two popular benchmark data sets VIPeR and CUHK01 validate the effectiveness of our proposed regularization approaches. Jianlou Si, Honggang Zhang 0002, Chun-Guang Li |
ICIP | 3 |
| 2015 | Improving tag matrix completion for image annotation and retrievalabstractImage annotation is a fundamental and challenging task in the field of semantic image retrieval. In this paper, we deal with image annotation via matrix completion. Concretely, we formulate the problem of annotating the tags of an image into a constrained optimization problem, in which the constraint is to keep the consistency with the given initial labels and the objective is to minimize the discrepancy between the correlation in visual content and the correlation in semantic tags. We solve the optimization problem with the linearized alternating direction method. Experimental results on benchmark data demonstrate the effectiveness of our proposals. Zhen Qin 0001, Chun-Guang Li, Honggang Zhang 0002, Jun Guo 0002 |
VCIP | 2 |
| 2014 | Person re-identification via region-of-interest based featuresabstractPerson re-identification is still a challenging task due to large visual appearance variations caused by illumination, background, viewpoints and poses in multi-camera surveillance. To address these challenges, many methods have been proposed. In this paper, we present an efficient method, called Region-of-Interest based Features (ROIF), via combining textural and chromatic features. It consists of two main phases - region-of-interest exploration from image and features extraction from ROI. Experimental results on the database VIPeR show that our method can yield promising accuracy with a quite cheap time cost. Jianlou Si, Honggang Zhang 0002, Chun-Guang Li |
VCIP | 3 |
| 2014 | Pairwise Rotation Invariant Co-Occurrence Local Binary PatternabstractDesigning effective features is a fundamental problem in computer vision. However, it is usually difficult to achieve a great tradeoff between discriminative power and robustness. Previous works shown that spatial co-occurrence can boost the discriminative power of features. However the current existing co-occurrence features are taking few considerations to the robustness and hence suffering from sensitivity to geometric and photometric variations. In this work, we study the Transform Invariance (TI) of co-occurrence features. Concretely we formally introduce a Pairwise Transform Invariance (PTI) principle, and then propose a novel Pairwise Rotation Invariant Co-occurrence Local Binary Pattern (PRICoLBP) feature, and further extend it to incorporate multi-scale, multi-orientation, and multi-channel information. Different from other LBP variants, PRICoLBP can not only capture the spatial context co-occurrence information effectively, but also possess rotation invariance. We evaluate PRICoLBP comprehensively on nine benchmark data sets from five different perspectives, e.g., encoding strategy, rotation invariance, the number of templates, speed, and discriminative power compared to other LBP variants. Furthermore we apply PRICoLBP to six different but related applications-texture, material, flower, leaf, food, and scene classification, and demonstrate that PRICoLBP is efficient, effective, and of a well-balanced tradeoff between the discriminative power and robustness. Xianbiao Qi, Rong Xiao 0003, Chun-Guang Li, Yu Qiao 0001, Jun Guo 0002, Xiaoou Tang |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2013 | Multi-scale Joint Encoding of Local Binary Patterns for Texture and Material Classification
Xianbiao Qi, Yu Qiao 0001, Chun-Guang Li, Jun Guo 0002 |
BMVC | 3 |
| 2013 | Exploring Cross-Channel Texture Correlation for Color Texture ClassificationabstractThis paper proposes a novel approach to encode cross-channel texture correlation for color texture classification task. Firstly, we quantitatively study the correlation between different color channels using Local Binary Pattern (LBP) as the texture descriptor and using Shannon’s information theory to measure the correlation. We find that (R, G) channel pair exhibits stronger correlation than (R, B) and (G, B) channel pairs. Secondly, we propose a novel descriptor to encode the cross-channel texture correlation. The proposed descriptor can capture well the relative variance of texture patterns between different channels. Meanwhile, our descriptor is computationally efficient and robust to image rotation. We conduct extensive experiments on four challenging color texture databases to validate the effectiveness of the proposed approach. The experimental results show that the proposed approach significantly outperforms its mostly relevant counterpart (Multichannel color LBP), and achieves the state-of-the-art performance. Xianbiao Qi, Yu Qiao 0001, Chun-Guang Li, Jun Guo 0002 |
BMVC | 3 |
| 2013 | Local alignment for query by hummingabstractQuery by humming (QBH) allows users to retrieve songs by humming a clip. In the previous work, the query has been regarded as a fragment of the music, so the task of QBH is considered to find a subsequence, which is most similar to the whole query, from the database. Taking into account humming errors, especially at the beginning or ending of the query, we assume that only part of the query is a subsequence of the music. Based on this assumption, we propose a local alignment framework which searches for the best match common subsequence between the query and database music. To verify the effectiveness of local alignment, two popular match algorithms, i.e. Linear Scaling and Dynamic Time Warping, are extended to identify the common subsequence. Experimental results on the 2010 MIREX-QBH corpus show that the new algorithms improve the retrieval accuracy significantly. Qiang Wang 0048, Gang Liu 0008, Chun-Guang Li, Jun Guo 0002 |
ICASSP | 4 |
| 2013 | Bases sorting: Generalizing the concept of frequency for over-complete dictionaries
Chun-Guang Li, Zhouchen Lin, Jun Guo 0002 |
Neurocomputing | 1 |
| 2012 | A rapid flower/leaf recognition systemabstractIn this work, we introduce a rapid and accurate flower/leaf recognition system. The system could process one query in less than 0.35s with users' simple interaction. Meanwhile, high accuracy and recall is achieved. Furthermore, low computational resource and memory cost are required by the system. Now, the system is demonstrated on 172 categories of flowers, the largest flower dataset until now, and 220 categories of leaves. Xianbiao Qi, Rong Xiao 0003, Lei Zhang 0001, Chun-Guang Li, Jun Guo 0002 |
ACM Multimedia | 4 |
| 2010 | Local Sparse Representation Based ClassificationabstractIn this paper, we address the computational complexity issue in Sparse Representation based Classification (SRC). In SRC, it is time consuming to find a global sparse representation. To remedy this deficiency, we propose a Local Sparse Representation based Classification (LSRC) scheme, which performs sparse decomposition in local neighborhood. In LSRC, instead of solving the l1-norm constrained least square problem for all of training samples we solve a similar problem in a local neighborhood for each test sample. Experiments on face recognition data sets ORL and Extended Yale B demonstrated that the proposed LSRC algorithm can reduce the computational complexity and remain the comparative classification accuracy and robustness. Chun-Guang Li, Jun Guo 0002, Honggang Zhang 0002 |
ICPR | 1 |
| 2009 | Learning Bundle Manifold by Double Neighborhood Graphs
Chun-Guang Li, Jun Guo 0002, Honggang Zhang 0002 |
ACCV (3) | 1 |
| 2009 | HCL2000 - A Large-scale Handwritten Chinese Character Database for Handwritten Character RecognitionabstractIn this paper, we present a large scale offline handwritten Chinese character database-HCL2000 which will be made public available for the research community. The database contains 3,755 frequently used simplified Chinese-characters written by 1,000 different subjects. The writerspsila information is incorporated in the database to facilitate testing on grouping writers with different background such as age, occupation, gender, and education etc. We investigate some characteristics of writing styles from different groups of writers. We evaluate HCL2000 database using three different algorithms as a baseline. We decide to publish the database along with this paper and make it free for a research purpose. Honggang Zhang 0002, Jun Guo 0002, Guang Chen 0003, Chun-Guang Li |
ICDAR | 4 |
| 2009 | Intrinsic Dimensionality Estimation within Neighborhood Convex HullabstractIn this paper, a novel method to estimate the intrinsic dimensionality of high-dimensional data set is proposed. Based on neighborhood information, our method calculates the non-negative locally linear reconstruction coefficients from its neighbors for each data point, and the numbers of those dominant positive reconstruction coefficients are regarded as a faithful guide to the intrinsic dimensionality of data set. The proposed method requires no parametric assumption on data distribution and is easy to implement in the general framework of manifold learning. Experimental results on several synthesized data sets and real data sets have shown the benefits of the proposed method. Chun-Guang Li, Jun Guo 0002 |
Int. J. Pattern Recognit. Artif. Intell. | 1 |