EDBT 2026 Demo / reviewers in the wild / expert
Guangcan Liu
dblp:07/3768 · also Guangchan Liu
· DBLP profile ↗
101ranked-venue papers
18as first author
34since 2021 · last 2026
0000-0002-9428-4387ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 53 · 13 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 53 · 5 first-author · 17 since 2021Databases, data management, data science and information retrieval · 8 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 since 2021Computer networks · 2 · 2 since 2021Theory of computation · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A negative-anchored self-relabeling strategy for multi-label class-incremental learning
Kaile Du, Junzhou Xie, Fan Lyu, Zihan Ye, Guangcan Liu |
Neurocomputing | 6 |
| 2026 | Negative-weighted knowledge distillation regularized graph convolutional network for multi-label class-incremental learning
Kaile Du, Junzhou Xie, Fan Lyu, Zihan Ye, Yuyang Li 0005, Guangcan Liu |
Pattern Recognit. | 8 |
| 2026 | Exact recovery of tensors with arbitrarily located corruptions through convolution principal component pursuit
Yuyang Li 0005, Kaile Du, Guangcan Liu |
Signal Process. | 4 |
| 2026 | Dynamic Prompting Spatial Temporal Actor Transformer for Fine-Grained Skeleton-Based Action RecognitionabstractThe scarcity of flexibility and effectiveness in skeleton models, combined with the characteristic of limited information in skeleton data, has resulted in fine-grained modeling insufficiently explored in recent skeleton-based action recognition. Confronting these challenges, we propose Dynamic prompting Spatial Temporal Actor transFormer (DSTAFormer), a powerful framework which neatly unifies vision and language. Specifically, we introduce a decoupled vision transformer, which consists of three components: Spatial transFormer (SF), Temporal transFormer (TF), and Actor transFormer (AF), to account for numerous visual aspects of the human body, namely spatial, temporal, and interactive relations. Compared to the vanilla transformer, we reformulate self-attention using Statistically-inspired Attention Reconstruction (SAR) module and Local-specific constraints, thereby enabling a more explicit and interpretable exploration of the action’s fine-grained compositions. The skeleton sequences are processed by this decoupled structure to generate the visual embeddings. To encode environmental interactions that skeletal coordinates inherently lack, we utilize Dynamic Prompting (Dp) strategy to generate visual-based textual prompts. These prompts are transformed into discriminative textual embeddings via a pre-trained large language model (LLM). We also design a Semantic Adapter (SA) to bridge the modality gap. The cross-modality embeddings are projected into a unified feature space for contrastive co-training. This infusion of knowledge into the skeleton data enhances its semantic richness, pushing the boundaries of fine-grained understanding. We evaluate our framework on NTU RGB+D, NTU RGB+D 120, and Toyota Smarthome datasets. DSTAFormer achieves comparable performance against state-of-the-arts. Yisheng Zhu, Guangcan Liu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | Breaking Low-Light Fusion Barrier: Unsupervised Darkness and Noise-Aware Visible and Infrared Image Fusion NetworkabstractInfrared and visible image fusion aims to integrate complementary information but suffers from severe residual degradations under low-light conditions. Existing methods face two main limitations: darkness-aware fusion focuses on illumination enhancement while neglecting noise suppression, and degradation-aware supervised fusion relies on paired data and struggles with enhancement-denoising imbalance due to multi-task optimization conflicts. To address these issues, we propose BLFusion, an unsupervised darkness- and noise-aware fusion framework that performs illumination enhancement and noise suppression in two dedicated stages while fusing complementary information without high-quality references. First, a Retinex-guided state space model-based decomposition network models illumination degradation to brighten dark visible images. Then, an unsupervised denoising fusion network jointly performs fusion and denoising, where noise correlation is disrupted by shuffling and a blind-spot network with dilated convolutions estimates clean representations from surrounding pixels. Finally, noise-free features from both modalities are fused to generate the final image. Moreover, we construct the MRLL dataset with 500 well-aligned infrared-visible image pairs, filling the gap for real-world noise-degraded nighttime scenarios. Experiments demonstrate that BLFusion outperforms state-of-the-art methods and generalizes robustly across diverse low-light and noisy conditions. The MRLL dataset and code are publicly available at https://github.com/ChenDoubleJ/BLFusion-MRLL. Han Xu 0001, Guangcan Liu, Jiayi Ma 0001 |
IEEE Trans. Image Process. | 3 |
| 2026 | Diff-MEF: Cross-Modal Diffusion Framework With Text Prompts and Semantic Perception for Multi-Exposure Image FusionabstractThe absence of real-world ground truth (GT) remains a challenge in multi-exposure image fusion (MEF). Benchmarks synthesizing pseudo GT through algorithm ensembles. Existing methods, hampered by inherent imperfections of pseudo GT and fixed mapping relationships, show limited performance and robustness. To address the limitations, we propose a novel cross-modal diffusion framework that synergizes text prompts and semantic perception for MEF, termed as Diff-MEF. First, it reformulates MEF as a probabilistic estimation task with conditional diffusion model for progressive transition and fusion. Then, we explicitly infer semantic and exposure priors as text prompts and semantic perception to improve performance and robustness. The priors are synergized through multi-modal prior embedding and optimization guidance. On the one hand, regarding cross-modal interaction, multi-modal priors, including segmentation masks, and exposure- and content-aware text prompts, are embedded into diffusion process by dedicated encoders and refine visual features through a text-segmentation refinement module. On the other hand, a semantic-level contrastive loss builds a regularization between cross-modal features in the semantic space of CLIP to mitigate degradations introduced by pseudo GT and fusion distortions. Experiments demonstrate that Diff-MEF outperforms SOTA methods and pseudo GT with superior fusion performance and robustness across diverse exposure scenarios. Code is available at https://github.com/hanna-xu/Diff-MEF. Han Xu 0001, Yunfei Huang, Linfeng Tang, Jiayi Ma 0001, Guangcan Liu |
IEEE Trans. Image Process. | 5 |
| 2025 | Rebalancing Multi-Label Class-Incremental LearningabstractMulti-label class-incremental learning (MLCIL) is essential for real-world multi-label applications, allowing models to learn new labels while retaining previously learned knowledge continuously. However, recent MLCIL approaches can only achieve suboptimal performance due to the oversight of the positive-negative imbalance problem, which manifests at both the label and loss levels because of the task-level partial label issue. The imbalance at the label level arises from the substantial absence of negative labels, while the imbalance at the loss level stems from the asymmetric contributions of the positive and negative loss parts to the optimization. To address the issue above, we propose a Rebalance framework for both the Loss and Label levels (RebLL), which integrates two key modules: asymmetric knowledge distillation (AKD) and online relabeling (OR). AKD is proposed to rebalance at the loss level by emphasizing the negative label learning in classification loss and down-weighting the contribution of overconfident predictions in distillation loss. OR is designed for label rebalance, which restores the original class distribution in memory by online relabeling the missing classes. Our comprehensive experiments on the PASCAL VOC and MS-COCO datasets demonstrate that this rebalancing strategy significantly improves performance, achieving new state-of-the-art results even with a vanilla CNN backbone. Kaile Du, Fan Lyu, Yuyang Li 0005, Junzhou Xie, Yixi Shen, Fuyuan Hu, Guangcan Liu |
AAAI | 8 |
| 2025 | Deno-IF: Unsupervised Noisy Visible and Infrared Image Fusion MethodabstractMost image fusion methods are designed for ideal scenarios and struggle to handle noise. Existing noise-aware fusion methods are supervised and heavily rely on constructed paired data, limiting performance and generalization. This paper proposes a novel unsupervised noisy visible and infrared image fusion method, comprising two key modules. First, when only noisy source images are available, a convolutional low-rank optimization module decomposes clean components based on convolutional low-rank priors, guiding subsequent optimization. The unsupervised approach eliminates data dependency and enhances generalization across various and variable noise. Second, a unified network jointly realizes denoising and fusion. It consists of both intra-modal recovery and inter-modal recovery and fusion, also with a convolutional low-rankness loss for regularization. By exploiting the commonalities of denoising and fusion, the joint framework significantly reduces network complexity while expanding functionality. Extensive experiments validate the effectiveness and generalization of the proposed method for image fusion under various and variable noise conditions. The code is publicly available at https://github.com/hanna-xu/Deno-IF. Han Xu 0001, Yuyang Li 0005, Yunfei Deng, Jiayi Ma 0001, Guangcan Liu |
NeurIPS | 5 |
| 2025 | URFusion: Unsupervised Unified Degradation-Robust Image Fusion NetworkabstractWhen dealing with low-quality source images, existing image fusion methods either fail to handle degradations or are restricted to specific degradations. This study proposes an unsupervised unified degradation-robust image fusion network, termed as URFusion, in which various types of degradations can be uniformly eliminated during the fusion process, leading to high-quality fused images. URFusion is composed of three core modules: intrinsic content extraction, intrinsic content fusion, and appearance representation learning and assignment. It first extracts degradation-free intrinsic content features from images affected by various degradations. These content features then provide feature-level rather than image-level fusion constraints for optimizing the fusion network, effectively eliminating degradation residues and reliance on ground truth. Finally, URFusion learns the appearance representation of images and assigns the statistical appearance representation of high-quality images to the content-fused result, producing the final high-quality fused image. Extensive experiments on multi-exposure image fusion and multi-modal image fusion tasks demonstrate the advantages of URFusion in fusion performance and suppression of multiple types of degradations. The code is available at https://github.com/hanna-xu/URFusion. Han Xu 0001, Xunpeng Yi, Chen Lu 0004, Guangcan Liu, Jiayi Ma 0001 |
IEEE Trans. Image Process. | 4 |
| 2024 | Confidence Self-calibration for Multi-label Class-Incremental Learning
Kaile Du, Fan Lyu, Yuyang Li 0005, Chen Lu 0004, Guangcan Liu |
ECCV (31) | 6 |
| 2024 | CTCFNet: CNN-Transformer Complementary and Fusion Network for High-Resolution Remote Sensing Image Semantic SegmentationabstractSemantic segmentation of high-resolution remote sensing images poses challenges such as scale variability, diverse objects, and obstruction by surface elements. These factors often lead existing methods to suffer from issues like missed and false detections, as well as coarse segmentation boundaries. To tackle these challenges, this article proposes a CNN-transformer complementary and fusion network, termed as CTCFNet. It aims to enhance segmentation accuracy and robustness by extracting and integrating the complementary global and local information from high-resolution remote sensing images. The CTCFNet operates through two primary stages: feature extraction and fusion. In the feature extraction stage, a feature extractor employs convolutional neural network (CNN) and pyramid vision transformer (PVT) blocks to extract both local and global features. A boundary loss is also proposed to improve the segmentation performance for object textures and boundaries. In the feature fusion stage, a feature aggregation module (FAM) is first designed to effectively fuse local and global features at the same scale, facilitating the feature extractor to obtain more comprehensive representations. On this basis, a bi-directional decoder (BiDecoder) reconstructs multiscale features through both top-down and bottom-up directions, resulting in more precise segmentation outputs. Experiments on several high-resolution remote sensing image datasets demonstrate that the proposed method outperforms the state-of-the-art methods in terms of segmentation accuracy and generalization. The code is available athttps://github.com/ChenLu0000/CTCFNet. Chen Lu 0004, Kaile Du, Han Xu 0001, Guangcan Liu |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Extraordinarily Time- and Memory-Efficient Large-Scale Canonical Correlation Analysis in Fourier Domain: From Shallow to DeepabstractCanonical correlation analysis (CCA) is a correlation analysis technique that is widely used in statistics and the machine-learning community. However, the high complexity involved in the training process lays a heavy burden on the processing units and memory system, making CCA nearly impractical in large-scale data. To overcome this issue, a novel CCA method that tries to carry out analysis on the dataset in the Fourier domain is developed in this article. Appling Fourier transform on the data, we can convert the traditional eigenvector computation of CCA into finding some predefined discriminative Fourier bases that can be learned with only element-wise dot product and sum operations, without complex time-consuming calculations. As the eigenvalues come from the sum of individual sample products, they can be estimated in parallel. Besides, thanks to the data characteristic of pattern repeatability, the eigenvalues can be well estimated with partial samples. Accordingly, a progressive estimate scheme is proposed, in which the eigenvalues are estimated through feeding data batch by batch until the eigenvalues sequence is stable in order. As a result, the proposed method shows its characteristics of extraordinarily fast and memory efficiencies. Furthermore, we extend this idea to the nonlinear kernel and deep models and obtained satisfactory accuracy and extremely fast training time consumption as expected. An extensive discussion on the fast Fourier transform (FFT)-CCA is made in terms of time and memory efficiencies. Experimental results on several large-scale correlation datasets, such as MNIST8M, X-RAY MICROBEAM SPEECH, and Twitter Users Data, demonstrate the superiority of the proposed algorithm over state-of-the-art (SOTA) large-scale CCA methods, as our proposed method achieves almost same accuracy with the training time of our proposed method being 1000 times faster. This makes our proposed models best practice models for dealing with large-scale correlation datasets. The source code is available at https://github.com/Mrxuzhao/FFTCCA. Xiangjun Shen, Zhaorui Xu, Liangjun Wang, Zechao Li, Guangcan Liu, Jianping Fan 0007, Zhengjun Zha |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | Decoupled spatio-temporal grouping transformer for skeleton-based action recognition
Shengkun Sun, Zihao Jia, Yisheng Zhu, Guangcan Liu, Zhengtao Yu 0001 |
Vis. Comput. | 4 |
| 2024 | A real-time semi-dense depth-guided depth completion network
JieJie Xu, Yisheng Zhu, Guangcan Liu |
Vis. Comput. | 4 |
| 2023 | Modeling the Relative Visual Tempo for Self-supervised Skeleton-based Action RecognitionabstractVisual tempo characterizes the dynamics and the temporal evolution, which helps describe actions. Recent approaches directly perform visual tempo prediction on skeleton sequences, which may suffer from insufficient feature representation issue. In this paper, we observe that relative visual tempo is more in line with human intuition, and thus providing more effective supervision signals. Based on this, we propose a novel Relative Visual Tempo Contrastive Learning framework for skeleton action Representation (RVTCLR). Specifically, we design a Relative Visual Tempo Learning (RVTL) task to explore the motion information in intra-video clips, and an Appearance-Consistency (AC) task to learn appearance information simultaneously, resulting in more representative spatiotemporal features. Furthermore, skeleton sequence data is much sparser than RGB data, making the network learn shortcuts, and overfit to low-level information such as skeleton scales. To learn high-order semantics, we further design a new Distribution-Consistency (DC) branch, containing three components: Skeleton-specific Data Augmentation (S-DA), Fine-grained Skeleton Encoding Module (FSEM), and Distribution-aware Diversity (DD) Loss. We term our entire method (RVTCLR with DC) as RVTCLR+. Extensive experiments on NTU RGB+D 60 and NTU RGB+D 120 datasets demonstrate that our RVTCLR+ can achieve competitive results over the state-of-the-art methods. Code is available at https://github.com/Zhuysheng/RVTCLR. Yisheng Zhu, Hu Han 0001, Zhengtao Yu 0001, Guangcan Liu |
ICCV | 4 |
| 2023 | Real-Time Traffic Sign Detection Based on Weighted Attention and Model Refinement
Zihao Jia, Shengkun Sun, Guangcan Liu |
Neural Process. Lett. | 3 |
| 2023 | Optimization Induced Equilibrium Networks: An Explicit Optimization Perspective for Understanding Equilibrium ModelsabstractTo reveal the mystery behind deep neural networks (DNNs), optimization may offer a good perspective. There are already some clues showing the strong connection between DNNs and optimization problems, e.g., under a mild condition, DNN's activation function is indeed a proximal operator. In this paper, we are committed to providing a unified optimization induced interpretability for a special class of networks-equilibrium models, i.e., neural networks defined by fixed point equations, which have become increasingly attractive recently. To this end, we first decompose DNNs into a new class of unit layer that is the proximal operator of an implicit convex function while keeping its output unchanged. Then, the equilibrium model of the unit layer can be derived, we name it Optimization Induced Equilibrium Networks (OptEq). The equilibrium point of OptEq can be theoretically connected to the solution of a convex optimization problem with explicit objectives. Based on this, we can flexibly introduce prior properties to the equilibrium points: 1) modifying the underlying convex problems explicitly so as to change the architectures of OptEq; and 2) merging the information into the fixed point iteration, which guarantees to choose the desired equilibrium point when the fixed point set is non-singleton. We show that OptEq outperforms previous implicit models even with fewer parameters. Xingyu Xie, Qiuhao Wang, Zenan Ling, Xia Li 0005, Guangcan Liu, Zhouchen Lin |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2023 | Sequential Order-Aware Coding-Based Robust Subspace Clustering for Human Action Recognition in Untrimmed VideosabstractHuman action recognition (HAR) is one of most important tasks in video analysis. Since video clips distributed on networks are usually untrimmed, it is required to accurately segment a given untrimmed video into a set of action segments for HAR. As an unsupervised temporal segmentation technology, subspace clustering learns the codes from each video to construct an affinity graph, and then cuts the affinity graph to cluster the video into a set of action segments. However, most of the existing subspace clustering schemes not only ignore the sequential information of frames in code learning, but also the negative effects of noises when cutting the affinity graph, which lead to inferior performance. To address these issues, we propose a sequential order-aware coding-based robust subspace clustering (SOAC-RSC) scheme for HAR. By feeding the motion features of video frames into multi-layer neural networks, two expressive code matrices are learned in a sequential order-aware manner from unconstrained and constrained videos, respectively, to construct the corresponding affinity graphs. Then, with the consideration of the existence of noise effects, a simple yet robust cutting algorithm is proposed to cut the constructed affinity graphs to accurately obtain the action segments for HAR. The extensive experiments demonstrate the proposed SOAC-RSC scheme achieves the state-of-the-art performance on the datasets of Keck Gesture and Weizmann, and provides competitive performance on the other 6 public datasets such as UCF101 and URADL for HAR task, compared to the recent related approaches. Zhili Zhou 0001, Chun Ding, Jin Li 0002, Eman Mohammadi, Guangcan Liu, Yimin Yang 0001, Q. M. Jonathan Wu |
IEEE Trans. Image Process. | 5 |
| 2023 | Multilevel Spatial-Temporal Excited Graph Network for Skeleton-Based Action RecognitionabstractThe ability to capture joint connections in complicated motion is essential for skeleton-based action recognition. However, earlier approaches may not be able to fully explore this connection in either the spatial or temporal dimension due to fixed or single-level topological structures and insufficient temporal modeling. In this paper, we propose a novel multilevel spatial-temporal excited graph network (ML-STGNet) to address the above problems. In the spatial configuration, we decouple the learning of the human skeleton into general and individual graphs by designing a multilevel graph convolution (ML-GCN) network and a spatial data-driven excitation (SDE) module, respectively. ML-GCN leverages joint-level, part-level, and body-level graphs to comprehensively model the hierarchical relations of a human body. Based on this, SDE is further introduced to handle the diverse joint relations of different samples in a data-dependent way. This decoupling approach not only increases the flexibility of the model for graph construction but also enables the generality to adapt to various data samples. In the temporal configuration, we apply the concept of temporal difference to the human skeleton and design an efficient temporal motion excitation (TME) module to highlight the motion-sensitive features. Furthermore, a simplified multiscale temporal convolution (MS-TCN) network is introduced to enrich the expression ability of temporal features. Extensive experiments on the four popular datasets NTU-RGB+D, NTU-RGB+D 120, Kinetics Skeleton 400, and Toyota Smarthome demonstrate that ML-STGNet gains considerable improvements over the existing state of the art. Yisheng Zhu, Hui Shuai, Guangcan Liu, Qingshan Liu 0001 |
IEEE Trans. Image Process. | 3 |
| 2023 | Recovery of Future Data via Convolution Nuclear Norm MinimizationabstractThis paper studies the problem of time series forecasting (TSF) from the perspective of compressed sensing. First of all, we convert TSF into a more inclusive problem called tensor completion with arbitrary sampling (TCAS), which is to restore a tensor from a subset of its entries sampled in an arbitrary manner. While it is known that, in the framework of Tucker low-rankness, it is theoretically impossible to identify the target tensor based on some arbitrarily selected entries, in this work we shall show that TCAS is indeed tackleable in the light of a new concept called convolutional low-rankness, which is a generalization of the well-known Fourier sparsity. Then we introduce a convex program termed Convolution Nuclear Norm Minimization (CNNM), and we prove that CNNM succeeds in solving TCAS as long as a sampling condition—which depends on the convolution rank of the target tensor—is obeyed. This theory provides a meaningful answer to the fundamental question of what is the minimum sampling size needed for making a given number of forecasts. Experiments on univariate time series, images and videos show encouraging results. Guangcan Liu, Wayne Zhang 0001 |
IEEE Trans. Inf. Theory | 1 |
| 2023 | Presswork defect inspection using only defect-free high-resolution images
Zhenyu Guan 0001, Yisheng Zhu, Guangcan Liu |
Vis. Comput. | 4 |
| 2023 | Multi-level feature fusion pyramid network for object detection
Zebin Guo, Hui Shuai, Guangcan Liu, Yisheng Zhu |
Vis. Comput. | 3 |
| 2022 | Orthogonal multi-view tensor-based learning for clustering
Shuangxun Ma, Yuehu Liu, Guangcan Liu, Qinghai Zheng, Chi Zhang 0020 |
Neurocomputing | 3 |
| 2022 | Self-Supervised Video Representation Learning Using Improved Instance-Wise Contrastive Learning and Deep ClusteringabstractInstance-wise contrastive learning (Instance-CL), which learns to map similar instances closer and different instances farther apart in the embedding space, has achieved considerable progress in self-supervised video representation learning. However, canonical Instance-CL does not handle properly the temporal similarities between different videos, limiting the representation capabilities of learned models. This paper presents a novel two-stage framework that combines Instance-CL and unsupervised clustering to progressively learn desirable temporal representations with high intra-class compactness. Specifically, (a) we first introduce a new consistency-preserving sampling strategy to generate positive/negative pairs. Compared to the traditional sampling methods, our sampling strategy focuses more on motion dynamics, resulting in more temporal-related feature representations. (b) To further explore the temporal similarities between videos so as to encourage intra-class compactness, we set temporal representations extracted from Instance-CL as an initializer, and iteratively use k-means clustering to generate pseudo-labels for training the encoder. We term our method as Improved Instance-CL with Deep Clustering (ICDC) and apply it to two downstream tasks, including action recognition and video retrieval. Extensive experimental results show that ICDC gains considerable improvements compared to the existing self-supervised methods. Yisheng Zhu, Hui Shuai, Guangcan Liu, Qingshan Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Predicting Tropical Cyclogenesis Using a Deep Learning Method From Gridded Satellite and ERA5 Reanalysis Data in the Western North Pacific BasinabstractThis article proposes a deep learning model to predict tropical cyclogenesis (TCG) from gridded satellite and ERA5 reanalysis data in the western North Pacific basin. The proposed model contains two modules. First, convolutional neural network (CNN)-based deep features are extracted for each predictor, and then, the extracted features are fused with two fully connected layers to differentiate and investigate the relationship between predictors and TCG. The experimental data of this study are composed of 3232 developing tropical cluster clouds and 6657 nondeveloping ones; 90% of the collected data are utilized to train the model, and the rest are used to evaluate the trained model. Totally, nine predictors have been considered for the study, and the results show that the brightness temperature (IR), relative vorticity (Vo), and geopotential height (Z) perform better than the other predictors. A combined model with six predictors [IR, Z, RH (relative humidity), Vo, WS10 m(wind speed at the height of ten meters above the surface of the Earth), and mslp (mean sea-level pressure)] achieves the best TCG predicting performance, i.e., 97.1% of developing tropical cyclones are detected at a probability threshold of 0.13 with a false alarm rate of 20.3%. The experimental results demonstrate that the proposed method is superior to the existing methods and also indicate that the fusion of satellite and reanalysis data is a promising method to predict TCG. Rui Zhang 0049, Qingshan Liu 0001, Renlong Hang, Guangcan Liu |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Time Series Forecasting via Learning Convolutionally Low-Rank ModelsabstractRecently, Liu and Zhang studied the rather challenging problem of time series forecasting from the perspective of compressed sensing. They proposed a no-learning method, named Convolution Nuclear Norm Minimization (CNNM), and proved that CNNM can exactly recover the future part of a series from its observed part, provided that the series is convolutionally low-rank. While impressive, the convolutional low-rankness condition may not be satisfied whenever the series is far from being seasonal, and is in fact brittle to the presence of trends and dynamics. This paper tries to approach the issues by integrating a learnable, orthonormal transformation into CNNM, with the purpose for converting the series of involute structures into regular signals of convolutionally low-rank. We prove that the resultant model, termed Learning-Based CNNM (LbCNNM), strictly succeeds in identifying the future part of a series, as long as the transform of the series is convolutionally low-rank. To learn proper transformations that may meet the required success conditions, we devise an interpretable method based on Principal Component Pursuit (PCP). Equipped with this learning method and some elaborate data argumentation skills, LbCNNM not only can handle well the major components of time series (including trends, seasonality and dynamics), but also can make use of the forecasts provided by some other forecasting methods; this means LbCNNM can be used as a general tool for model combination. Extensive experiments on 100,452 real-world time series from Time Series Data Library (TSDL) and M4 Competition (M4) demonstrate the superior performance of LbCNNM. Guangcan Liu |
IEEE Trans. Inf. Theory | 1 |
| 2021 | Tensor LISTA: Differentiable sparse representation learning for multi-dimensional tensor
Qi Zhao 0013, Guangcan Liu, Qingshan Liu 0001 |
Neurocomputing | 2 |
| 2021 | Matrix Completion with Deterministic Sampling: Theories and MethodsabstractIn some significant applications such as data forecasting, the locations of missing entries cannot obey any non-degenerate distributions, questioning the validity of the prevalent assumption that the missing data is randomly chosen according to some probabilistic model. To break through the limits of random sampling, we explore in this paper the problem of real-valued matrix completion under the setup of deterministic sampling. We propose two conditions, isomeric condition and relative well-conditionedness, for guaranteeing an arbitrary matrix to be recoverable from a sampling of the matrix entries. It is provable that the proposed conditions are weaker than the assumption of uniform sampling and, most importantly, it is also provable that the isomeric condition is necessary for the completions of any partial matrices to be identifiable. Equipped with these new tools, we prove a collection of theorems for missing data recovery as well as convex/nonconvex matrix completion. Among other things, we study in detail a Schatten quasi-norm induced method termed isomeric dictionary pursuit (IsoDP), and we show that IsoDP exhibits some distinct behaviors absent in the traditional bilinear programs. Guangcan Liu, Qingshan Liu 0001, Xiao-Tong Yuan, Meng Wang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2021 | Parallel Connected LSTM for Matrix Sequence Prediction with Elusive CorrelationsabstractThis article is about a challenging problem called matrix sequence prediction, which is motivated from the application of taxi order prediction. Remarkably, the problem differs greatly from previous sequence prediction tasks in the sense that the time-wise correlations are quite elusive; namely, distant entries could be strongly correlated and nearby entries are unnecessarily related. Such distinct specifics make prevalent convolution-recurrence-based methods inadequate to apply. To remedy this trouble, we propose a novel architecture called Parallel Connected LSTM (PcLSTM), which integrates two new mechanisms, Multi-channel Linearized Connection (McLC) and Adaptive Parallel Unit (APU), into the framework of LSTM. Benefiting from the strengths of McLC and APU, our PcLSTM is able to handle well both the elusive correlations within each timestamp and the temporal dependencies across different timestamps, achieving state-of-the-art performance in a set of experiments demonstrated on synthetic and real-world datasets. Qi Zhao 0013, Chuqiao Chen, Guangcan Liu, Qingshan Liu 0001, Shengyong Chen |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2021 | Collaborative Local-Global Learning for Temporal Action ProposalabstractTemporal action proposal generation is an essential and challenging task in video understanding, which aims to locate the temporal intervals that likely contain the actions of interest. Although great progress has been made, the problem is still far from being well solved. In particular, prevalent methods can handle well only the local dependencies (i.e., short-term dependencies) among adjacent frames but are generally powerless in dealing with the global dependencies (i.e., long-term dependencies) between distant frames. To tackle this issue, we propose CLGNet, a novel Collaborative Local-Global Learning Network for temporal action proposal. The majority of CLGNet is an integration of Temporal Convolution Network and Bidirectional Long Short-Term Memory, in which Temporal Convolution Network is responsible for local dependencies while Bidirectional Long Short-Term Memory takes charge of handling the global dependencies. Furthermore, an attention mechanism called the background suppression module is designed to guide our model to focus more on the actions. Extensive experiments on two benchmark datasets, THUMOS’14 and ActivityNet-1.3, show that the proposed method can outperform state-of-the-art methods, demonstrating the strong capability of modeling the actions with varying temporal durations. Yisheng Zhu, Hu Han 0001, Guangcan Liu, Qingshan Liu 0001 |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2021 | Flexible Auto-Weighted Local-Coordinate Concept Factorization: A Robust Framework for Unsupervised ClusteringabstractConcept Factorization (CF) and its variants may produce inaccurate representation and clustering results due to the sensitivity to noise, hard constraint on the reconstruction error, and pre-obtained approximate similarities. To improve the representation ability, a novel unsupervised Robust Flexible Auto-weighted Local-coordinate Concept Factorization (RFA-LCF) framework is proposed for clustering high-dimensional data. Specifically, RFA-LCF integrates the robust flexible CF by clean data space recovery, robust sparse local-coordinate coding, and adaptive weighting into a unified model. RFA-LCF improves the representations by enhancing the robustness of CF to noise and errors, providing a flexible constraint on the reconstruction error and optimizing the locality jointly. For robust learning, RFA-LCF clearly learns a sparse projection to recover the underlying clean data space, and then the flexible CF is performed in the projected feature space. RFA-LCF also uses a L2,1-norm based flexible residue to encode the mismatch between the recovered data and its reconstruction, and uses the robust sparse local-coordinate coding to represent data using a few nearby basis concepts. For auto-weighting, RFA-LCF jointly preserves the manifold structures in the basis concept space and new coordinate space in an adaptive manner by minimizing the reconstruction errors on clean data, anchor points and coordinates. By updating the local-coordinate preserving data, basis concepts and new coordinates alternately, the representation abilities can be potentially improved. Extensive results on public databases show that RFA-LCF delivers the state-of-the-art clustering results compared with other related methods. Zhao Zhang 0001, Yan Zhang 0053, Sheng Li 0001, Guangcan Liu, Dan Zeng 0001, Shuicheng Yan, Meng Wang 0001 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2021 | Twin-Incoherent Self-Expressive Locality-Adaptive Latent Dictionary Pair Learning for ClassificationabstractThe projective dictionary pair learning (DPL) model jointly seeks a synthesis dictionary and an analysis dictionary by extracting the block-diagonal coefficients with an incoherence-constrained analysis dictionary. However, DPL fails to discover the underlying subspaces and salient features at the same time, and it cannot encode the neighborhood information of the embedded coding coefficients, especially adaptively. In addition, although the data can be well reconstructed via the minimization of the reconstruction error, useful distinguishing salient feature information may be lost and incorporated into the noise term. In this article, we propose a novel self-expressive adaptive locality-preserving framework: twin-incoherent self-expressive latent DPL (SLatDPL). To capture the salient features from the samples, SLatDPL minimizes a latent reconstruction error by integrating the coefficient learning and salient feature extraction into a unified model, which can also be used to simultaneously discover the underlying subspaces and salient features. To make the coefficients block diagonal and ensure that the salient features are discriminative, our SLatDPL regularizes them by imposing a twin-incoherence constraint. Moreover, SLatDPL utilizes a self-expressive adaptive weighting strategy that uses normalized block-diagonal coefficients to preserve the locality of the codes and salient features. SLatDPL can use the class-specific reconstruction residual to handle new data directly. Extensive simulations on several public databases demonstrate the satisfactory performance of our SLatDPL compared with related methods. Zhao Zhang 0001, Yang Wang 0023, Zheng Zhang 0006, Haijun Zhang 0002, Guangcan Liu, Meng Wang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2021 | 3D Tensor Auto-encoder with Application to Video CompressionabstractAuto-encoder has been widely used to compress high-dimensional data such as the images and videos. However, the traditional auto-encoder network needs to store a large number of parameters. Namely, when the input data is of dimension n , the number of parameters in an auto-encoder is in general O ( n ). In this article, we introduce a network structure called 3D Tensor Auto-Encoder (3DTAE). Unlike the traditional auto-encoder, in which a video is represented as a vector, our 3DTAE considers videos as 3D tensors to directly pass tensor objects through the network. The weights of each layer are represented by three small matrices, and thus the number of parameters in 3DTAE is just O ( n 1/3). The compact nature of 3DTAE fits well the needs of video compression. Given an ensemble of high-dimensional videos, we represent them as 3DTAE networks plus some small core tensors, and we further quantize the network parameters and the core tensors to get the final compressed data. Experimental results verify the efficiency of 3DTAE. Yang Li 0039, Guangcan Liu, Yubao Sun, Qingshan Liu 0001, Shengyong Chen |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2021 | DLRF-Net: A Progressive Deep Latent Low-Rank Fusion Network for Hierarchical Subspace DiscoveryabstractLow-rank coding-based representation learning is powerful for discovering and recovering the subspace structures in data, which has obtained an impressive performance; however, it still cannot obtain deep hidden information due to the essence of single-layer structures. In this article, we investigate the deep low-rank representation of images in a progressive way by presenting a novel strategy that can extend existing single-layer latent low-rank models into multiple layers. Technically, we propose a new progressive Deep Latent Low-Rank Fusion Network (DLRF-Net) to uncover deep features and the clustering structures embedded in latent subspaces. The basic idea of DLRF-Net is to progressively refine the principal and salient features in each layer from previous layers by fusing the clustering and projective subspaces, respectively, which can potentially learn more accurate features and subspaces. To obtain deep hidden information, DLRF-Net inputs shallow features from the last layer into subsequent layers. Then, it aims at recovering the hierarchical information and deeper features by respectively congregating the subspaces in each layer of the network. As such, one can also ensure the representation learning of deeper layers to remove the noise and discover the underlying clean subspaces, which will be verified by simulations. It is noteworthy that the framework of our DLRF-Net is general and is applicable to most existing latent low-rank representation models, i.e., existing latent low-rank models can be easily extended to the multilayer scenario using DLRF-Net. Extensive results on real databases show that our framework can deliver enhanced performance over other related techniques. Zhao Zhang 0001, Jiahuan Ren, Haijun Zhang 0002, Zheng Zhang 0006, Guangcan Liu, Shuicheng Yan |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2020 | Multilayer Collaborative Low-Rank Coding Network for Robust Deep Subspace DiscoveryabstractFor subspace recovery, most existing low-rank representation (LRR) models performs in the original space in single-layer mode. As such, the deep hierarchical information cannot be learned, which may result in inaccurate recoveries for complex real data. In this paper, we explore the deep multi-subspace recovery problem by designing a multilayer architecture for latent LRR. Technically, we propose a new Multilayer Collaborative Low-Rank Representation Network model termed DeepLRR to discover deep features and deep subspaces. In each layer (>2), DeepLRR bilinearly reconstructs the data matrix by the collaborative representation with low-rank coefficients and projection matrices in the previous layer. The bilinear low-rank reconstruction of previous layer is directly fed into the next layer as the input and low-rank dictionary for representation learning, and is further decomposed into a deep principal feature part, a deep salient feature part and a deep sparse error. As such, the coherence issue can be also resolved due to the low-rank dictionary, and the robustness against noise can also be enhanced in the feature subspace. To recover the sparse errors in layers accurately, a dynamic growing strategy is used, as the noise level will become smaller for the increase of layers. Besides, a neighborhood reconstruction error is also included to encode the locality of deep salient features by deep coefficients adaptively in each layer. Extensive results on public databases show that our DeepLRR outperforms other related models for subspace discovery and clustering. Xianzhen Li, Zhao Zhang 0001, Yang Wang 0023, Guangcan Liu, Shuicheng Yan, Meng Wang 0001 |
ECAI | 4 |
| 2020 | Maximum-and-Concatenation NetworksabstractWhile successful in many fields, deep neural networks (DNNs) still suffer from some open problems such as bad local minima and unsatisfactory generalization performance. In this work, we propose a novel architecture called Maximum-and-Concatenation Networks (MCN) to try eliminating bad local minima and improving generalization ability as well. Remarkably, we prove that MCN has a very nice property; that is, every local minimum of an (l+1)-layer MCN can be better than, at least as good as, the global minima of the network consisting of its first l layers. In other words, by increasing the network depth, MCN can autonomously improve its local minima’s goodness, what is more, it is easy to plug MCN into an existing deep model to make it also have this property. Finally, under mild conditions, we show that MCN can approximate certain continuous function arbitrarily well with high efficiency; that is, the covering number of MCN is much smaller than most existing DNNs such as deep ReLU. Based on this, we further provide a tight generalization bound to guarantee the inference ability of MCN when dealing with testing samples. Xingyu Xie, Hao Kong 0002, Jianlong Wu, Wayne Zhang 0001, Guangcan Liu, Zhouchen Lin |
ICML | 5 |
| 2020 | Deep Latent Low-Rank Fusion Network for Progressive Subspace DiscoveryabstractLow-rank representation is powerful for recover-ing and clustering the subspace structures, but it cannot obtain deep hierarchical information due to the single-layer mode. In this paper, we present a new and effective strategy to extend the sin-gle-layer latent low-rank models into multi-ple-layers, and propose a new and progressive Deep Latent Low-Rank Fusion Network (DLRF-Net) to uncover deep features and struc-tures embedded in input data. The basic idea of DLRF-Net is to refine features progressively from the previous layers by fusing the subspaces in each layer, which can potentially obtain accurate fea-tures and subspaces for representation. To learn deep information, DLRF-Net inputs shallow fea-tures of the last layers into subsequent layers. Then, it recovers the deeper features and hierar-chical information by congregating the projective subspaces and clustering subspaces respectively in each layer. Thus, one can learn hierarchical sub-spaces, remove noise and discover the underlying clean subspaces. Note that most existing latent low-rank coding models can be extended to multi-layers using DLRF-Net. Extensive results show that our network can deliver enhanced perfor-mance over other related frameworks. Zhao Zhang 0001, Jiahuan Ren, Zheng Zhang 0006, Guangcan Liu |
IJCAI | 4 |
| 2020 | Learning image compressed sensing with sub-pixel convolutional generative adversarial network
Yubao Sun, Qingshan Liu 0001, Guangcan Liu |
Pattern Recognit. | 4 |
| 2020 | Selecting Optimal Completion to Partial Matrix via Self-ValidationabstractIn many applications such as film recommendation, one often encounters the problem of estimating the unseen entries in a partially observed matrix, formally known as matrix completion. Over the past several decades, lots of effective methods have been established in the literature, and each method may contain several hyper-parameters. For a partial matrix, one can use those methods with certain parametric settings to obtain a large number of completions. Now, a critical question is, how to select the optimal completion from a number of candidates? This question is indeed a hard to answer, because in practice the true values of the missing entries are unknown. Thus far, the only approach for dealing with the issue is through data-validation, which is to first split the observations into two subsets, a training set and a validation set, and then choose the model that performs best on the validation set as the winner to produce the final results. Though straightforward, this approach might fall in a non-optimal model that overfits the validation set. In this work, we shall suggest a different approach called self-validation, which accounts on a special metric that can evaluate the “goodness” of a completion without using any validation data. The metric is derived from the recently established isomeric condition, measuring the identifiable degree of the completion itself. Extensive experiments demonstrate that our self-validation approach is better than the commonly used data-validation. Guangcan Liu, Wayne Zhang 0001, Qingshan Liu 0001 |
IEEE Signal Process. Lett. | 2 |
| 2020 | Learning Hybrid Representation by Robust Dictionary Learning in Factorized Compressed SpaceabstractIn this paper, we investigate the robust dictionary learning (DL) to discover the hybrid salient low-rank and sparse representation in a factorized compressed space. A Joint Robust Factorization and Projective Dictionary Learning (J-RFDL) model is presented. The setting of J-RFDL aims at improving the data representations by enhancing the robustness to outliers and noise in data, encoding the reconstruction error more accurately and obtaining hybrid salient coefficients with accurate reconstruction ability. Specifically, J-RFDL performs the robust representation by DL in a factorized compressed space to eliminate the negative effects of noise and outliers on the results, which can also make the DL process efficient. To make the encoding process robust to noise in data, J-RFDL clearly uses sparse L2, 1-norm that can potentially minimize the factorization and reconstruction errors jointly by forcing rows of the reconstruction errors to be zeros. To deliver salient coefficients with good structures to reconstruct given data well, J-RFDL imposes the joint low-rank and sparse constraints on the embedded coefficients with a synthesis dictionary. Based on the hybrid salient coefficients, we also extend J-RFDL for the joint classification and propose a discriminative J-RFDL model, which can improve the discriminating abilities of learnt coefficients by minimizing the classification error jointly. Extensive experiments on public datasets demonstrate that our formulations can deliver superior performance over other state-of-the-art methods. Jiahuan Ren, Zhao Zhang 0001, Sheng Li 0001, Yang Wang 0023, Guangcan Liu, Shuicheng Yan, Meng Wang 0001 |
IEEE Trans. Image Process. | 5 |
| 2020 | Joint Label Prediction Based Semi-Supervised Adaptive Concept Factorization for Robust Data RepresentationabstractConstrained Concept Factorization (CCF) yields the enhanced representation ability over CF by incorporating label information as additional constraints, but it cannot classify and group unlabeled data appropriately. Minimizing the difference between the original data and its reconstruction directly can enable CCF to model a small noisy perturbation, but is not robust to gross sparse errors. Besides, CCF cannot preserve the manifold structures in new representation space explicitly, especially in an adaptive manner. In this paper, we propose a joint label prediction based Robust Semi-Supervised Adaptive Concept Factorization (RS2ACF) framework. To obtain robust representation, RS2ACF relaxes the factorization to make it simultaneously stable to small entrywise noise and robust to sparse errors. To enrich prior knowledge to enhance the discrimination, RS2ACF clearly uses class information of labeled data and more importantly propagates it to unlabeled data by jointly learning an explicit label indicator for unlabeled data. By the label indicator, RS2ACF can ensure the unlabeled data of the same predicted label to be mapped into the same class in feature space. Besides, RS2ACF incorporates the joint neighborhood reconstruction error over the new representations and predicted labels of both labeled and unlabeled data, so the manifold structures can be preserved explicitly and adaptively in the representation space and label space at the same time. Owing to the adaptive manner, the tricky process of determining the neighborhood size or kernel width can be avoided. Extensive results on public databases verify that our RS2ACF can deliver state-of-the-art data representation, compared with other related methods. Zhao Zhang 0001, Yan Zhang 0053, Guangcan Liu, Jinhui Tang 0001, Shuicheng Yan, Meng Wang 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2020 | Fine-grained action recognition using multi-view attentions
Yisheng Zhu, Guangcan Liu |
Vis. Comput. | 2 |
| 2019 | Robust Unsupervised Flexible Auto-weighted Local-coordinate Concept Factorization for Image ClusteringabstractWe investigate the high-dimensional data clustering problem by proposing a novel and unsupervised representation learning model called Robust Flexible Auto-weighted Local-coordinate Concept Factorization (RFA-LCF). RFA-LCF integrates the robust flexible CF, robust sparse local-coordinate coding and the adaptive reconstruction weighting learning into a unified model. The adaptive weighting is driven by including thejoint manifold preserving constraints on the recovered clean data, basis concepts and new representation. Specifically, our RFA-LCF uses a L2,1-norm based flexible residue to encode the mismatch between clean data and its reconstruction, and also applies the robust adaptive sparse local-coordinate coding to represent the data using a few nearby basis concepts, which can make the factorization more accurate and robust to noise. The robust flexible factorization is also performed in the recovered clean data space for enhancing representations. RFA-LCF also considers preserving the local manifold structures of clean data space, basis concept space and the new coordinate space jointly in an adaptive manner way. Extensive comparisons show that RFA-LCF can deliver enhanced clustering results. Zhao Zhang 0001, Yan Zhang 0053, Sheng Li 0001, Guangcan Liu, Meng Wang 0001, Shuicheng Yan |
ICASSP | 4 |
| 2019 | Learning Structured Twin-Incoherent Twin-Projective Latent Dictionary Pairs for ClassificationabstractIn this paper, we extend the popular dictionary pair learning (DPL) into the scenario of twin-projective latent flexible DPL under a structured twin-incoherence. Technically, a novel framework called Twin-Projective Latent Flexible DPL (TP-DPL) is proposed, which minimizes the twin-incoherence constrained flexibly-relaxed reconstruction error to avoid the possible over-fitting issue and produce accurate reconstruction. In this setting, TP-DPL integrates the twin-incoherence based latent flexible DPL and the joint embedding of codes as well as salient features by twin-projection into a unified model in an adaptive neighborhood-preserving manner. Therefore, TP-DPL can unify the procedures of salient feature representation and classification. The twin-incoherence constraint on coefficients and features can explicitly ensure high intra-class compactness and inter-class separation over them. TP-DPL also integrates the adaptive weighting to preserve local neighborhood of both coefficients and salient features within each class explicitly. For efficiency, TP-DPL selects the Frobenius-norm and abandons the costly l0/l1-norm for group sparse representation. Another byproduct is that TP-DPL can directly apply the class-specific twin-projective reconstruction residual to compute the label of data. Extensive results on public databases show that TP-DPL can deliver the state-of-the-art performance. Zhao Zhang 0001, Zheng Zhang 0006, Yang Wang 0023, Guangcan Liu, Meng Wang 0001 |
ICDM | 5 |
| 2019 | Differentiable Linearized ADMMabstractRecently, a number of learning-based optimization methods that combine data-driven architectures with the classical optimization algorithms have been proposed and explored, showing superior empirical performance in solving various ill-posed inverse problems, but there is still a scarcity of rigorous analysis about the convergence behaviors of learning-based optimization. In particular, most existing analyses are specific to unconstrained problems but cannot apply to the more general cases where some variables of interest are subject to certain constraints. In this paper, we propose Differentiable Linearized ADMM (D-LADMM) for solving the problems with linear constraints. Specifically, D-LADMM is a K-layer LADMM inspired deep neural network, which is obtained by firstly introducing some learnable weights in the classical Linearized ADMM algorithm and then generalizing the proximal operator to some learnable activation function. Notably, we rigorously prove that there exist a set of learnable parameters for D-LADMM to generate globally converged solutions, and we show that those desired parameters can be attained by training D-LADMM in a proper way. To the best of our knowledge, we are the first to provide the convergence analysis for the learning-based optimization method on constrained problems. Xingyu Xie, Jianlong Wu, Guangcan Liu, Zhisheng Zhong, Zhouchen Lin |
ICML | 3 |
| 2019 | Scalable Block-Diagonal Locality-Constrained Projective Dictionary LearningabstractWe propose a novel structured discriminative block- diagonal dictionary learning method, referred to as scalable Locality-Constrained Projective Dictionary Learning (LC-PDL), for efficient representation and classification. To improve the scalability by saving both training and testing time, our LC-PDL aims at learning a structured discriminative dictionary and a block-diagonal representation without using costly l0/l1-norm. Besides, it avoids extra time-consuming sparse reconstruction process with the well-trained dictionary for new sample as many existing models. More importantly, LC-PDL avoids using the com- plementary data matrix to learn the sub-dictionary over each class. To enhance the performance, we incorporate a locality constraint of atoms into the DL procedures to keep local information and obtain the codes of samples over each class separately. A block-diagonal discriminative approximation term is also derived to learn a discriminative projection to bridge data with their codes by extracting the special block-diagonal features from data, which can ensure the approximate coefficients to associate with its label information clearly. Then, a robust multiclass classifier is trained over extracted block-diagonal codes for accurate label predictions. Experimental results verify the effectiveness of our algorithm. Zhao Zhang 0001, Weiming Jiang, Zheng Zhang 0006, Sheng Li 0001, Guangcan Liu, Jie Qin 0004 |
IJCAI | 5 |
| 2019 | Moving object detection via segmentation and saliency constrained RPCA
Yang Li 0039, Guangcan Liu, Qingshan Liu 0001, Yubao Sun, Shengyong Chen |
Neurocomputing | 2 |
| 2019 | Matrix recovery with implicitly low-rank data
Xingyu Xie, Jianlong Wu, Guangcan Liu, Jun Wang 0039 |
Neurocomputing | 3 |
| 2019 | Robust auto-weighted projective low-rank and sparse recovery for visual representation
Lei Wang 0124, Bangjun Wang, Zhao Zhang 0001, Qiaolin Ye, Liyong Fu, Guangcan Liu, Meng Wang 0001 |
Neural Networks | 6 |
| 2019 | Robust Low-rank subspace segmentation with finite mixture noise
Xianglin Guo, Xingyu Xie, Guangcan Liu, Mingqiang Wei, Jun Wang 0039 |
Pattern Recognit. | 3 |
| 2019 | Kernel-Induced Label Propagation by Mapping for Semi-Supervised ClassificationabstractKernel methods have been successfully applied to the areas of pattern recognition and data mining. In this paper, we mainly discuss the issue of propagating labels in kernel space. A Kernel-Induced Label Propagation (Kernel-LP) framework by mapping is proposed for high-dimensional data classification using the most informative patterns of data in kernel space. The essence of Kernel-LP is to perform joint label propagation and adaptive weight learning in a transformed kernel space. That is, our Kernel-LP changes the task of label propagation from the commonly-used Euclidean space in most existing work to kernel space. The motivation of our Kernel-LP to propagate labels and learn the adaptive weights jointly by the assumption of an inner product space of inputs, i.e., the original linearly inseparable inputs may be mapped to be separable in kernel space. Kernel-LP is based on existing positive and negative LP model, i.e., the effects of negative label information are integrated to improve the label prediction power. Also, Kernel-LP performs adaptive weight construction over the same kernel space, so it can avoid the tricky process of choosing the optimal neighborhood size suffered in traditional criteria. Two novel and efficient out-of-sample approaches for our Kernel-LP to involve new test data are also presented, i.e., (1) direct kernel mapping and (2) kernel mapping-induced label reconstruction, both of which purely depend on the kernel matrix between training set and testing set. Owing to the kernel trick, our algorithms will be applicable to handle the high-dimensional real data. Extensive results on real datasets demonstrate the effectiveness of our approach. Zhao Zhang 0001, Lei Jia 0002, Ming-Bo Zhao, Guangcan Liu, Meng Wang 0001, Shuicheng Yan |
IEEE Trans. Big Data | 4 |
| 2019 | Robust Subspace Clustering With Compressed DataabstractDimension reduction is widely regarded as an effective way for decreasing the computation, storage and communication loads of data-driven intelligent systems, leading to a growing demand for statistical methods that allow analysis (e.g., clustering) of compressed data. We therefore study in this paper a novel problem called compressive robust subspace clustering, which is to perform robust subspace clustering with the compressed data, and which is generated by projecting the original high-dimensional data onto a lower-dimensional subspace chosen at random. Given only the compressed data and sensing matrix, the proposed method, row space pursuit (RSP), recovers the authentic row space that gives correct clustering results under certain conditions. Extensive experiments show that RSP is distinctly better than the competing methods, in terms of both clustering accuracy and computational efficiency. Guangcan Liu, Zhao Zhang 0001, Qingshan Liu 0001, Hongkai Xiong |
IEEE Trans. Image Process. | 1 |
| 2019 | Unsupervised Nonnegative Adaptive Feature Extraction for Data RepresentationabstractIn this paper, we propose a novel unsupervised Nonnegative Adaptive Feature Extraction (NAFE) algorithm for data representation and classification. The formulation of NAFE integrates the sparsity constrained nonnegative matrix factorization (NMF), representation learning, and adaptive reconstruction weight learning into a unified model. Specifically, NAFE performs feature and weight learning over the new robust representations of NMF for more accurate measure and representation. For nonnegative adaptive feature extraction, our NAFE first utilizes the sparsity constrained NMF to obtain the new and robust representations of the original data. To preserve the manifold structures of the learnt new representations, we also incorporate a neighborhood reconstruction error over the weight matrix for joint minimization. Note that to further improve the representation power, the weights are jointly shared in the new low-dimensional nonnegative representation space, low-dimensional nonlinear manifold space, and low-dimensional projective subspace, i.e., local neighborhood information is clearly preserved in different feature spaces so that informative representations and features can be jointly obtained. To enable NAFE to extract features from new data, we also include a feature approximation error by a linear projection so that the learnt extractor can obtain features from new data efficiently. Extensive simulations show that our formulation can deliver state-of-the-art results on several public databases for feature extraction and classification, compared with several related methods. Yan Zhang 0053, Zhao Zhang 0001, Sheng Li 0001, Jie Qin 0004, Guangcan Liu, Meng Wang 0001, Shuicheng Yan |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2019 | On fusing the latent deep CNN feature for image classification
Xueliang Liu, Rongjie Zhang, Richang Hong, Guangcan Liu |
World Wide Web | 5 |
| 2019 | Correction to: On fusing the latent deep CNN feature for image classification
Xueliang Liu, Rongjie Zhang, Richang Hong, Guangcan Liu |
World Wide Web | 5 |
| 2018 | Modeling Attention and Memory for Auditory Selection in a Cocktail Party EnvironmentabstractDeveloping a computational auditory model to solve the cocktail party problem has long bedeviled scientists, especially for a single microphone recording. Although recent deep learning based frameworks have made significant progress in multi-talker mixed speech separation, most existing deep learning based methods, focusing on separating all the speech channels rather than selectively attending the target speech and ignoring other sounds, may fail to offer a satisfactory solution in a complex auditory scene where the number of input sounds is usually uncertain and even dynamic. In this work, we employ ideas from auditory selective attention of behavioral and cognitive neurosciences and from recent advances of memory-augmented neural networks. Specifically, a unified Auditory Selection framework with Attention and Memory (dubbed ASAM) is proposed. Our ASAM first accumulates the prior knowledge (that is the acoustic feature to one specific speaker) into a life-long memory during the training phase, meanwhile a speech perceptor is trained to extract the temporal acoustic feature and update the memory online when a salient speech is given. Then, the learned memory is utilized to interact with the mixture input to attend and filter the target frequency out from the mixture stream. Finally, the network is trained to minimize the reconstruction error of the attended speech. We evaluate the proposed approach on WSJ0 and THCHS-30 datasets and the experimental results demonstrate that our approach successfully conducts two auditory selection tasks: the top-down task-specific attention (e.g. to follow a conversation with friend) and the bottom-up stimulus-driven attention (e.g. be attracted by a salient speech). Compared with deep clustering based methods, our method conducts competitive advantages especially in a real noise environment (e.g. street junction). Our code is available at https://github.com/jacoxu/ASAM. Jiaming Xu 0001, Jing Shi 0003, Guangcan Liu, Xiuyi Chen, Bo Xu 0002 |
AAAI | 3 |
| 2018 | Robust Projective Low-Rank and Sparse Representation by Robust Dictionary LearningabstractIn this paper, we discuss the robust factorization based robust dictionary learning problem for data representation. A Robust Projective Low-Rank and Sparse Representation model (R-PLSR) is technically proposed. Our R-PLSR model integrates the L1-norm based robust factorization and robust low-rank & sparse representation by robust dictionary learning into a unified framework. Specifically, R-PLSR performs the joint low-rank and sparse representation over the informative low-dimensional representations by robust sparse factorization so that the results are more accurate. To make the factorization and representation procedures robust to noise and outliers, R-PLSR imposes the sparse L2, 1-norm jointly on the reconstruction errors based on the factorization and dictionary learning. Note that L2, 1-norm can also minimize the reconstruction error as much as possible, since the L2, 1-norm theoretically tends to force many rows of the reconstruction error matrix to be zeros. The Nuclear-norm and L1-norm are jointly used on the representation coefficients so that salient representations can be obtained. Extensive results on several image datasets show that our R-PLSR formulation can deliver superior performance over other state-of-the-arts. Jiahuan Ren, Zhao Zhang 0001, Sheng Li 0001, Guangcan Liu, Meng Wang 0001, Shuicheng Yan |
ICPR | 4 |
| 2018 | Robust Discriminative Projective Dictionary Pair Learning by Adaptive RepresentationsabstractIn this paper, we mainly propose a Robust Adaptive Projective Dictionary Pair Learning (RA-DPL) framework based on the adaptive discriminative representations. Our formulation can seamlessly integrate the robust projective dictionary pair learning and the adaptive sparse representation learning into a unified model. RA-DPL improves the existing DPL algorithm in threefold. First, RA-DPL aims at computing the robust projective dictionary pairs by employing the sparse and robust l2,1-norm to encode the reconstruction error. Second, RA-DPL regularizes the robust l2,1-norm on the analysis dictionary so that the analysis dictionary can extract sparse coefficients from the given samples explicitly. More importantly, the optimization of l2,1-norm is so efficient, that is, the sparse coding step will be time-saving. Third, RA-DPL can clearly preserve the local neighborhood relationship of the sparse coefficients within each class, which can make the learnt representations discriminating and can also improve the discriminating power of learnt dictionary. Extensive simulations on image databases demonstrate that our RA-DPL can obtain the superior performance over other state-of-the-arts. Zhao Zhang 0001, Weiming Jiang, Guangcan Liu, Meng Wang 0001, Shuicheng Yan |
ICPR | 4 |
| 2018 | Robust Adaptive Low-Rank and Sparse Embedding for Feature RepresentationabstractMost existing low-rank sparse embedding models extract features of data in the original input space and usually separate the manifold preservation step from the coding process, which may result in the decreased performance. In this paper, a novel Robust Adaptive Low-rank and Sparse Embedding (RALSE) framework is technically proposed for salient feature extraction of the high-dimensional data by seamlessly integrating the joint low-rank and sparse recovery with the robust adaptive salient feature extraction. Specifically, our RALSE integrates the joint low-rank and sparse representation, adaptive neighborhood preserving graph weight learning and the robustness-promoting representation into a unified framework. For accurate similarity measure, RALSE computes the adaptive weights by minimizing the reconstruction error over the noise-removed data and salient features simultaneously, where L1-norm is regularized to ensure the sparse properties of learnt weights. RALSE can also ensure the learnt projection to preserve local neighborhood information of embedded features clearly and adaptively. The projection is not only modeled under joint low-rank and sparse regularization, but also computed from a clean subspace, making it powerful for the salient feature extraction. Thus, the learnt low-rank sparse features would be more accurate for subsequent classification. Extensive results demonstrate the effectiveness of our RALSE formulation for data representation and classification. Lei Wang 0124, Zhao Zhang 0001, Guangcan Liu, Qiaolin Ye, Jie Qin 0004, Meng Wang 0001 |
ICPR | 3 |
| 2018 | Robust Locality-Constrained Label Consistent K-SVD by Joint Sparse EmbeddingabstractWe mainly propose a robust Embedded Locality-Constrained Label Consistent Dictionary Learning (ELC2DL) framework for discriminative classification. ELC2DL improves the representation and classification performance by performing DL in the noise-removed sparse embedding space, since most real data often contains noise and performing DL over noisy data for reconstruction may decrease performance potentially. To reduce the noise in data, our model computes a sparse projection jointly for noise reduction and then uses the noise-removed data for DL. By incorporating a noise-reduction term with a discriminative locality-constrained label consistent term that associates the label information with each dictionary atom to preserve local structure of training data, a noise-reduction projection, an over-complete dictionary and discriminative sparse codes are obtained jointly. Simulations on several image databases show that our algorithm can deliver enhanced performance over other state-of-the-arts. Zhao Zhang 0001, Weiming Jiang, Sheng Li 0001, Jie Qin 0004, Guangcan Liu, Shuicheng Yan |
ICPR | 5 |
| 2018 | Listen, Think and Listen Again: Capturing Top-down Auditory Attention for Speaker-independent Speech SeparationabstractRecent deep learning methods have made significant progress in multi-talker mixed speech separation. However, most existing models adopt a driftless strategy to separate all the speech channels rather than selectively attend the target one. As a result, those frameworks may be failed to offer a satisfactory solution in complex auditory scene where the number of input sounds is usually uncertain and even dynamic. In this paper, we present a novel neural network based structure motivated by the top-down attention behavior of human when facing complicated acoustical scene. Different from previous works, our method constructs an inference-attention structure to predict interested candidates and extract each speech channel of them. Our work gets rid of the limitation that the number of channels must be given or the high computation complexity for label permutation problem. We evaluated our model on the WSJ0 mixed-speech tasks. In all the experiments, our model gets highly competitive to reach and even outperform the baselines. Jing Shi 0003, Jiaming Xu 0001, Guangcan Liu, Bo Xu 0002 |
IJCAI | 3 |
| 2018 | Distilled Binary Neural Network for Monaural Speech SeparationabstractMonaural speech separation, aiming at solving the cocktail party problem, has many important application scenarios, most of which ask for the real-time response, high energy efficiency and efficient storage. However, the state-of-the-art Deep Neural Network based separation models usually require huge memory and computation for the 32-bit floating point multiply accumulations, hence most of them cannot meet those requirements. Recently, there are many methods proposed to solve the problem, and binary neural networks have drawn many attentions for they compress and speed up its counterparts at the cost of some performance. Hence, in this paper, we binarize Deep Neural Network based separation models, aiming to deploy them on embedded devices for real-time applications. Furthermore, we improve the separation performance by integrating knowledge distillation into the training phase of binary neural network based models, which is referred as Distilled Binary Neural Network (DBNN). To the best of our knowledge, DBNN is the first attempt to integrate two types of model compression. In the experiments, we demonstrate the effectiveness of our proposed method, which successfully binarizes the Deep Neural Network based separation models with a comparable performance. Xiuyi Chen, Guangcan Liu, Jing Shi 0003, Jiaming Xu 0001, Bo Xu 0002 |
IJCNN | 2 |
| 2018 | Improving Speech Separation with Adversarial Network and Reinforcement LearningabstractIn contrast to the conventional deep neural network for single-channel speech separation, we propose a separation framework based on adversarial network and reinforcement learning. The purpose of the adversarial network inspired by the generative adversarial network is to make the separated result and ground-truth with the same data distribution by evaluating the discrepancy between them. Meanwhile, in order to enable the model to bias the generation towards desirable metrics and reduce the discrepancy between training loss (such as mean squared error) and testing metric (such as SDR), we present the future success based on reinforcement learning. We directly optimize the performance metric to accomplish exactly that. With the combination of adversarial network and reinforcement learning, our model is able to improve the performance of single-channel speech separation. Guangcan Liu, Jing Shi 0003, Xiuyi Chen, Jiaming Xu 0001, Bo Xu 0002 |
IJCNN | 1 |
| 2018 | Similarity-Adaptive Latent Low-Rank Representation for Robust Data Representation
Lei Wang 0124, Zhao Zhang 0001, Sheng Li 0001, Guangcan Liu, Chenping Hou, Jie Qin 0004 |
PRICAI (1) | 4 |
| 2018 | Multi-scale Spatiotemporal Information Fusion Network for Video Action RecognitionabstractTwo-stream convolutional networks have shown excellent performance in video action recognition in recent years. However, it remains unclear how to model the correlation between the temporal and spatial streams more effectively. First, the spatial stream and temporal stream pay attention to different aspects, which can lead to different recognition results. Second, the variety in the length of optical flow fields tends to have a great impact on the classification results. In this paper, we propose a novel multi-scale spatiotemporal information fusion network to fuse the spatial and temporal features. Specifically, our network takes advantage of multi-scale temporal information to better utilize the motion cues. Considering the complementary relationship between the spatial and temporal features, we take the hierarchical fusion strategies and asynchronous fusion method to fuse the two-stream features. Experimental results on two benchmark datasets (UCF101 and HMDB51) show that the proposed network achieves competitive performance. Yutong Cai, Weiyao Lin, John See, Ming-Ming Cheng, Guangcan Liu, Hongkai Xiong |
VCIP | 5 |
| 2018 | Multi-component group sparse RPCA model for motion object detection under complex dynamic background
Yubao Sun, Renlong Hang, Qingshan Liu 0001, Guangcan Liu |
Neurocomputing | 5 |
| 2018 | Spatio-temporal convolutional features with nested LSTM for facial expression recognition
Zhenbo Yu, Guangcan Liu, Qingshan Liu 0001, Jiankang Deng |
Neurocomputing | 2 |
| 2018 | Semi-supervised vehicle classification via fusing affinity matrices
Maojin Sun, Shijie Hao, Guangcan Liu |
Signal Process. | 3 |
| 2018 | Implicit Block Diagonal Low-Rank RepresentationabstractWhile current block diagonal constrained subspace clustering methods are performed explicitly on the original data space, in practice, it is often more desirable to embed the block diagonal prior into the reproducing kernel Hilbert feature space by kernelization techniques, as the underlying data structure in reality is usually nonlinear. However, it is still unknown how to carry out the embedding and kernelization in the models with block diagonal constraints. In this paper, we shall take a step in this direction. First, we establish a novel model termed implicit block diagonal low-rank representation (IBDLR), by incorporating the implicit feature representation and block diagonal prior into the prevalent low-rank representation method. Second, mostly important, we show that the model in IBDLR could be kernelized by making use of a smoothed dual representation and the specifics of a proximal gradient-based optimization algorithm. Finally, we provide some theoretical analyses for the convergence of our optimization algorithm. Comprehensive experiments on synthetic and real-world data sets demonstrate the superiorities of our IBDLR over state-of-the-art methods. Xingyu Xie, Xianglin Guo, Guangcan Liu, Jun Wang 0039 |
IEEE Trans. Image Process. | 3 |
| 2018 | Reversed Spectral HashingabstractHashing is emerging as a powerful tool for building highly efficient indices in large-scale search systems. In this paper, we study spectral hashing (SH), which is a classical method of unsupervised hashing. In general, SH solves for the hash codes by minimizing an objective function that tries to preserve the similarity structure of the data given. Although computationally simple, very often SH performs unsatisfactorily and lags distinctly behind the state-of-the-art methods. We observe that the inferior performance of SH is mainly due to its imperfect formulation; that is, the optimization of the minimization problem in SH actually cannot ensure that the similarity structure of the high-dimensional data is really preserved in the low-dimensional hash code space. In this paper, we, therefore, introduce reversed SH (ReSH), which is SH with its input and output interchanged. Unlike SH, which estimates the similarity structure from the given high-dimensional data, our ReSH defines the similarities between data points according to the unknown low-dimensional hash codes. Equipped with such a reversal mechanism, ReSH can seamlessly overcome the drawback of SH. More precisely, the minimization problem in our ReSH can be optimized if and only if similar data points are mapped to adjacent hash codes, and mostly important, dissimilar data points are considerably separated from each other in the code space. Finally, we solve the minimization problem in ReSH by multilayer neural networks and obtain state-of-the-art retrieval results on three benchmark data sets. Qingshan Liu 0001, Guangcan Liu, Lai Li, Xiao-Tong Yuan, Meng Wang 0001, Wei Liu 0005 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2018 | Deeper cascaded peak-piloted network for weak expression recognition
Zhenbo Yu, Qinshan Liu, Guangcan Liu |
Vis. Comput. | 3 |
| 2017 | A New Theory for Matrix CompletionabstractPrevalent matrix completion theories reply on an assumption that the locations of the missing data are distributed uniformly and randomly (i.e., uniform sampling). Nevertheless, the reason for observations being missing often depends on the unseen observations themselves, and thus the missing data in practice usually occurs in a nonuniform and deterministic fashion rather than randomly. To break through the limits of random sampling, this paper introduces a new hypothesis called \emph{isomeric condition}, which is provably weaker than the assumption of uniform sampling and arguably holds even when the missing data is placed irregularly. Equipped with this new tool, we prove a series of theorems for missing data recovery and matrix completion. In particular, we prove that the exact solutions that identify the target matrix are included as critical points by the commonly used nonconvex programs. Unlike the existing theories for nonconvex matrix completion, which are built upon the same condition as convex programs, our theory shows that nonconvex programs have the potential to work with a much weaker condition. Comparing to the existing studies on nonuniform sampling, our setup is more general. Guangcan Liu, Qingshan Liu 0001, Xiao-Tong Yuan |
NIPS | 1 |
| 2017 | Blessing of Dimensionality: Recovering Mixture Data via Dictionary PursuitabstractThis paper studies the problem of recovering the authentic samples that lie on a union of multiple subspaces from their corrupted observations. Due to the high-dimensional and massive nature of today's data-driven community, it is arguable that the target matrix (i.e., authentic sample matrix) to recover is often low-rank. In this case, the recently established Robust Principal Component Analysis (RPCA) method already provides us a convenient way to solve the problem of recovering mixture data. However, in general, RPCA is not good enough because the incoherent condition assumed by RPCA is not so consistent with the mixture structure of multiple subspaces. Namely, when the subspace number grows, the row-coherence of data keeps heightening and, accordingly, RPCA degrades. To overcome the challenges arising from mixture data, we suggest to consider LRR in this paper. We elucidate that LRR can well handle mixture data, as long as its dictionary is configured appropriately. More precisely, we mathematically prove that LRR can weaken the dependence on the row-coherence, provided that the dictionary is well-conditioned and has a rank of not too high. In particular, if the dictionary itself is sufficiently low-rank, then the dependence on the row-coherence can be completely removed. These provide some elementary principles for dictionary learning and naturally lead to a practical algorithm for recovering mixture data. Our experiments on randomly generated matrices and real motion sequences show promising results. Guangcan Liu, Qingshan Liu 0001, Ping Li 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2017 | Adaptive Cascade Regression Model For Robust Face AlignmentabstractCascade regression is a popular face alignment approach, and it has achieved good performances on the wild databases. However, it depends heavily on local features in estimating reliable landmark locations and therefore suffers from corrupted images, such as images with occlusion, which often exists in real-world face images. In this paper, we present a new adaptive cascade regression model for robust face alignment. In each iteration, the shape-indexed appearance is introduced to estimate the occlusion level of each landmark, and each landmark is then weighted according to its estimated occlusion level. Also, the occlusion levels of the landmarks act as adaptive weights on the shape-indexed features to decrease the noise on the shape-indexed features. At the same time, an exemplar-based shape prior is designed to suppress the influence of local image corruption. Extensive experiments are conducted on the challenging benchmarks, and the experimental results demonstrate that the proposed method achieves better results than the state-of-the-art methods for facial landmark localization and occlusion detection. Qingshan Liu 0001, Jiankang Deng, Jing Yang 0038, Guangcan Liu, Dacheng Tao |
IEEE Trans. Image Process. | 4 |
| 2016 | Advancing Iterative Quantization Hashing Using Isotropic Prior
Lai Li, Guangcan Liu, Qingshan Liu 0001 |
MMM (2) | 2 |
| 2016 | Learning Additive Exponential Family Graphical Models via \ell_{2, 1}-norm Regularized M-EstimationabstractWe investigate a subclass of exponential family graphical models of which the sufficient statistics are defined by arbitrary additive forms. We propose two $\ell_{2,1}$-norm regularized maximum likelihood estimators to learn the model parameters from i.i.d. samples. The first one is a joint MLE estimator which estimates all the parameters simultaneously. The second one is a node-wise conditional MLE estimator which estimates the parameters for each node individually. For both estimators, statistical analysis shows that under mild conditions the extra flexibility gained by the additive exponential family models comes at almost no cost of statistical efficiency. A Monte-Carlo approximation method is developed to efficiently optimize the proposed estimators. The advantages of our estimators over Gaussian graphical models and Nonparanormal estimators are demonstrated on synthetic and real data sets. Xiao-Tong Yuan, Ping Li 0001, Tong Zhang 0001, Qingshan Liu 0001, Guangcan Liu |
NIPS | 5 |
| 2016 | A Deterministic Analysis for LRRabstractThe recently proposed low-rank representation (LRR) method has been empirically shown to be useful in various tasks such as motion segmentation, image segmentation, saliency detection and face recognition. While potentially powerful, LRR depends heavily on the configuration of its key parameter, λ. In realistic environments where the prior knowledge about data is lacking, however, it is still unknown how to choose λ in a suitable way. Even more, there is a lack of rigorous analysis about the success conditions of the method, and thus the significance of LRR is a little bit vague. In this paper we therefore establish a theoretical analysis for LRR, striving for figuring out under which conditions LRR can be successful, and deriving a moderately good estimate to the key parameter λ as well. Simulations on synthetic data points and experiments on real motion sequences verify our claims. Guangcan Liu, Huan Xu 0001, Jinhui Tang 0001, Qingshan Liu 0001, Shuicheng Yan |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2016 | Multitask Low-Rank Affinity Graph for Image Segmentation and Image AnnotationabstractThis article investigates a low-rank representation--based graph, which can used in graph-based vision tasks including image segmentation and image annotation. It naturally fuses multiple types of image features in a framework named multitask low-rank affinity pursuit. Given the image patches described with multiple types of features, we aim at inferring a unified affinity matrix that implicitly encodes the relations among these patches. This is achieved by seeking the sparsity-consistent low-rank affinities from the joint decompositions of multiple feature matrices into pairs of sparse and low-rank matrices, the latter of which is expressed as the production of the image feature matrix and its corresponding image affinity matrix. The inference process is formulated as a minimization problem and solved efficiently with the augmented Lagrange multiplier method. Considering image patches as vertices, a graph can be built based on the resulted affinity matrix. Compared to previous methods, which are usually based on a single type of feature, the proposed method seamlessly integrates multiple types of features to jointly produce the affinity matrix in a single inference step. The proposed method is applied to graph-based image segmentation and graph-based image annotation. Experiments on benchmark datasets well validate the superiority of using multiple features over single feature and also the superiority of our method over conventional methods for feature fusion. Teng Li 0001, Bin Cheng 0001, Bingbing Ni, Guangcan Liu, Shuicheng Yan |
ACM Trans. Intell. Syst. Technol. | 4 |
| 2015 | Low-Rank Tensor Constrained Multiview Subspace ClusteringabstractIn this paper, we explore the problem of multiview subspace clustering. We introduce a low-rank tensor constraint to explore the complementary information from multiple views and, accordingly, establish a novel method called Low-rank Tensor constrained Multiview Subspace Clustering (LT-MSC). Our method regards the subspace representation matrices of different views as a tensor, which captures dexterously the high order correlations underlying multiview data. Then the tensor is equipped with a low-rank constraint, which models elegantly the cross information among different views, reduces effectually the redundancy of the learned subspace representations, and improves the accuracy of clustering as well. The inference process of the affinity matrix for clustering is formulated as a tensor nuclear norm minimization problem, constrained with an additional L2,1-norm regularizer and some linear equalities. The minimization problem is convex and thus can be solved efficiently by an Augmented Lagrangian Alternating Direction Minimization (AL-ADM) method. Extensive experimental results on four benchmark datasets show the effectiveness of our proposed LT-MSC method. Changqing Zhang 0002, Huazhu Fu, Si Liu 0001, Guangcan Liu, Xiaochun Cao |
ICCV | 4 |
| 2015 | Non-blind deblurring of structured images with geometric deformation
Xin Zhang 0051, Fuchun Sun 0001, Guangcan Liu, Yi Ma 0001 |
Vis. Comput. | 3 |
| 2014 | A Latent Clothing Attribute Approach for Human Pose Estimation
Jie Shen 0005, Guangcan Liu, Yong Yu 0001 |
ACCV (1) | 3 |
| 2014 | Recovery of Coherent Data via Low-Rank Dictionary Pursuit
Guangcan Liu, Ping Li 0001 |
NIPS | 1 |
| 2014 | Blind Image Deblurring Using Spectral Properties of Convolution OperatorsabstractBlind deconvolution is to recover a sharp version of a given blurry image or signal when the blur kernel is unknown. Because this problem is ill-conditioned in nature, effectual criteria pertaining to both the sharp image and blur kernel are required to constrain the space of candidate solutions. While the problem has been extensively studied for long, it is still unclear how to regularize the blur kernel in an elegant, effective fashion. In this paper, we show that the blurry image itself actually encodes rich information about the blur kernel, and such information can indeed be found by exploring and utilizing a well-known phenomenon, that is, sharp images are often high pass, whereas blurry images are usually low pass. More precisely, we shall show that the blur kernel can be retrieved through analyzing and comparing how the spectrum of an image as a convolution operator changes before and after blurring. Subsequently, we establish a convex kernel regularizer, which depends only on the given blurry image. Interestingly, the minimizer of this regularizer guarantees to give a good estimate to the desired blur kernel if the original image is sharp enough. By combining this powerful regularizer with the prevalent nonblind devonvolution techniques, we show how we could significantly improve the deblurring results through simulations on synthetic images and experiments on realistic images. Guangcan Liu, Shiyu Chang, Yi Ma 0001 |
IEEE Trans. Image Process. | 1 |
| 2014 | Unified Structured Learning for Simultaneous Human Pose Estimation and Garment Attribute ClassificationabstractIn this paper, we utilize structured learning to simultaneously address two intertwined problems: 1) human pose estimation (HPE) and 2) garment attribute classification (GAC), which are valuable for a variety of computer vision and multimedia applications. Unlike previous works that usually handle the two problems separately, our approach aims to produce an optimal joint estimation for both HPE and GAC via a unified inference procedure. To this end, we adopt a preprocessing step to detect potential human parts from each image (i.e., a set of candidates) that allows us to have a manageable input space. In this way, the simultaneous inference of HPE and GAC is converted to a structured learning problem, where the inputs are the collections of candidate ensembles, outputs are the joint labels of human parts and garment attributes, and joint feature representation involves various cues such as pose-specific features, garment-specific features, and cross-task features that encode correlations between human parts and garment attributes. Furthermore, we explore the strong edge evidence around the potential human parts so as to derive more powerful representations for oriented human parts. Such evidences can be seamlessly integrated into our structured learning model as a kind of energy function, and the learning process could be performed by standard structured support vector machines algorithm. However, the joint structure of the two problems is a cyclic graph, which hinders efficient inference. To resolve this issue, we compute instead approximate optima using an iterative procedure, where in each iteration, the variables of one problem are fixed. In this way, satisfactory solutions can be efficiently computed by dynamic programming. Experimental results on two benchmark data sets show the state-of-the-art performance of our approach. Jie Shen 0005, Guangcan Liu, Jia Chen 0001, Yuqiang Fang, Jianbin Xie, Yong Yu 0001, Shuicheng Yan |
IEEE Trans. Image Process. | 2 |
| 2014 | Fast Low-Rank Subspace SegmentationabstractSubspace segmentation is the problem of segmenting (or grouping) a set of$n$data points into a number of clusters, with each cluster being a (linear) subspace. The recently established algorithms such as Sparse Subspace Clustering (SSC), Low-Rank Representation (LRR) and Low-Rank Subspace Segmentation (LRSS) are effective in terms of segmentation accuracy, but computationally inefficient as they possess a complexity of$O(n^{3})$, which is too high to afford for the case where$n$is very large. In this paper we devise a fast subspace segmentation algorithm with complexity of$O(n\log (n))$. This is achieved by firstly using partial Singular Value Decomposition (SVD) to approximate the solution of LRSS, secondly utilizing Locality Sensitive Hashing (LSH) to build a sparse affinity graph that encodes the subspace memberships, and finally adopting a fast Normalized Cut (NCut) algorithm to produce the final segmentation results. Besides of high efficiency, our algorithm also has comparable effectiveness as the original LRSS method. Xin Zhang 0051, Fuchun Sun 0001, Guangcan Liu, Yi Ma 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2013 | Robust Recovery of Subspace Structures by Low-Rank RepresentationabstractIn this paper, we address the subspace clustering problem. Given a set of data samples (vectors) approximately drawn from a union of multiple subspaces, our goal is to cluster the samples into their respective subspaces and remove possible outliers as well. To this end, we propose a novel objective function named Low-Rank Representation (LRR), which seeks the lowest rank representation among all the candidates that can represent the data samples as linear combinations of the bases in a given dictionary. It is shown that the convex program associated with LRR solves the subspace clustering problem in the following sense: When the data is clean, we prove that LRR exactly recovers the true subspace structures; when the data are contaminated by outliers, we prove that under certain conditions LRR can exactly recover the row space of the original data and detect the outlier as well; for data corrupted by arbitrary sparse errors, LRR can also approximately recover the row space with theoretical guarantees. Since the subspace membership is provably determined by the row space, these further imply that LRR can perform robust subspace clustering and error correction in an efficient and effective way. Guangcan Liu, Zhouchen Lin, Shuicheng Yan, Ju Sun, Yong Yu 0001, Yi Ma 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2013 | Improving Bottom-up Saliency Detection by Looking into NeighborsabstractBottom-up saliency detection aims to detect salient areas within natural images usually without learning from labeled images. Typically, the saliency map of an image is inferred by only using the information within this image (referred to as the “current image”). While efficient, such single-image-based methods may fail to obtain reliable results, because the information within a single image may be insufficient for defining saliency. In this paper, we investigate how saliency detection can benefit from the nearest neighbor structure in the image space. First, we show that existing methods can be improved by extending them to include the visual neighborhood information. This verifies the significance of the neighbors. Next, a solution of multitask sparsity pursuit is proposed to integrate the current image and its neighbors to collaboratively detect saliency. The integration is done by first representing each image as a feature matrix, and then seeking the consistently sparse elements from the joint decompositions of multiple matrices into pairs of low-rank and sparse matrices. The computational procedure is formulated as a constrained nuclear norm and ℓ2,1-norm minimization problem, which is convex and can be solved efficiently with the augmented Lagrange multiplier method. Besides the nearest neighbor structure in the visual feature space, the proposed model can also be generalized to handle multiple visual features. Extensive experiments have clearly validated its superiority over other state-of-the-art methods. Congyan Lang, Jiashi Feng, Guangcan Liu, Jinhui Tang 0001, Shuicheng Yan, Jiebo Luo 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2013 | General Subspace Learning With Corrupted Training Data Via Graph EmbeddingabstractWe address the following subspace learning problem: supposing we are given a set of labeled, corrupted training data points, how to learn the underlying subspace, which contains three components: an intrinsic subspace that captures certain desired properties of a data set, a penalty subspace that fits the undesired properties of the data, and an error container that models the gross corruptions possibly existing in the data. Given a set of data points, these three components can be learned by solving a nuclear norm regularized optimization problem, which is convex and can be efficiently solved in polynomial time. Using the method as a tool, we propose a new discriminant analysis (i.e., supervised subspace learning) algorithm called Corruptions Tolerant Discriminant Analysis (CTDA), in which the intrinsic subspace is used to capture the features with high within-class similarity, the penalty subspace takes the role of modeling the undesired features with high between-class similarity, and the error container takes charge of fitting the possible corruptions in the data. We show that CTDA can well handle the gross corruptions possibly existing in the training data, whereas previous linear discriminant analysis algorithms arguably fail in such a setting. Extensive experiments conducted on two benchmark human face data sets and one object recognition data set show that CTDA outperforms the related algorithms. Bing-Kun Bao, Guangcan Liu, Richang Hong, Shuicheng Yan, Changsheng Xu |
IEEE Trans. Image Process. | 2 |
| 2012 | Street-to-shop: Cross-scenario clothing retrieval via parts alignment and auxiliary setabstractIn this paper, we address a practical problem of cross-scenario clothing retrieval - given a daily human photo captured in general environment, e.g., on street, finding similar clothing in online shops, where the photos are captured more professionally and with clean background. There are large discrepancies between daily photo scenario and online shopping scenario. We first propose to alleviate the human pose discrepancy by locating 30 human parts detected by a well trained human detector. Then, founded on part features, we propose a two-step calculation to obtain more reliable one-to-many similarities between the query daily photo and online shopping photos: 1) the within-scenario one-to-many similarities between a query daily photo and the auxiliary set are derived by direct sparse reconstruction; and 2) by a cross-scenario many-to-many similarity transfer matrix inferred offline from an extra auxiliary set and the online shopping set, the reliable cross-scenario one-to-many similarities between the query daily photo and all online shopping photos are obtained. We collect a large online shopping dataset and a daily photo dataset, both of which are thoroughly labeled with 15 clothing attributes via Mechanic Turk. The extensive experimental evaluations on the collected datasets well demonstrate the effectiveness of the proposed framework for cross-scenario clothing retrieval. Si Liu 0001, Guangcan Liu, Changsheng Xu, Hanqing Lu, Shuicheng Yan |
CVPR | 3 |
| 2012 | Practical low-rank matrix approximation under robust L1-normabstractA great variety of computer vision tasks, such as rigid/nonrigid structure from motion and photometric stereo, can be unified into the problem of approximating a low-rank data matrix in the presence of missing data and outliers. To improve robustness, the L1-norm measurement has long been recommended. Unfortunately, existing methods usually fail to minimize the L1-based nonconvex objective function sufficiently. In this work, we propose to add a convex trace-norm regularization term to improve convergence, without introducing too much heterogenous information. We also customize a scalable first-order optimization algorithm to solve the regularized formulation on the basis of the augmented Lagrange multiplier (ALM) method. Extensive experimental results verify that our regularized formulation is reasonable, and the solving algorithm is very efficient, insensitive to initialization and robust to high percentage of missing data and/or outliers1. Yinqiang Zheng, Guangcan Liu, Shigeki Sugimoto, Shuicheng Yan, Masatoshi Okutomi |
CVPR | 2 |
| 2012 | Multi-Task low-rank and sparse matrix recovery for human motion segmentationabstractThis paper proposes a new algorithm, named Multi-Task Robust Principal Component Analysis (MTRPCA), to collaboratively integrate multiple visual features and motion priors for human motion segmentation. Given the video data described by multiple features, the human motion part is obtained by jointly decomposing multiple feature matrices into pairs of low-rank and sparse matrices. The inference process is formulated as a convex optimization problem that minimizes a constrained combination of nuclear norm and ℓ2,1-norm, which can be solved efficiently with Augmented Lagrange Multiplier (ALM) method. Compared to previous methods, which usually make use of individual features, the proposed method seamlessly integrates multiple features and priors within a single inference step, and thus produces more accurate and reliable results. Experiments on the HumanEva human motion dataset show that the proposed MTRPCA is novel and promising. Xiangyang Wang 0003, Guangcan Liu |
ICIP | 3 |
| 2012 | Active Subspace: Toward Scalable Low-Rank LearningabstractWe address the scalability issues in low-rank matrix learning problems. Usually these problems resort to solving nuclear norm regularized optimization problems (NNROPs), which often suffer from high computational complexities if based on existing solvers, especially in large-scale settings. Based on the fact that the optimal solution matrix to an NNROP is often low rank, we revisit the classic mechanism of low-rank matrix factorization, based on which we present an active subspace algorithm for efficiently solving NNROPs by transforming large-scale NNROPs into small-scale problems. The transformation is achieved by factorizing the large solution matrix into the product of a small orthonormal matrix (active subspace) and another small matrix. Although such a transformation generally leads to nonconvex problems, we show that a suboptimal solution can be found by the augmented Lagrange alternating direction method. For the robust PCA (RPCA) (Candès, Li, Ma, & Wright, 2009 ) problem, a typical example of NNROPs, theoretical results verify the suboptimality of the solution produced by our algorithm. For the general NNROPs, we empirically show that our algorithm significantly reduces the computational complexity without loss of optimality. Guangcan Liu, Shuicheng Yan |
Neural Comput. | 1 |
| 2012 | Inductive Robust Principal Component AnalysisabstractIn this paper we address the error correction problem that is to uncover the low-dimensional subspace structure from high-dimensional observations, which are possibly corrupted by errors. When the errors are of Gaussian distribution, Principal Component Analysis (PCA) can find the optimal (in terms of least-square-error) low-rank approximation to highdimensional data. However, the canonical PCA method is known to be extremely fragile to the presence of gross corruptions. Recently, Wright et al. established a so-called Robust Principal Component Analysis (RPCA) method, which can well handle grossly corrupted data [14]. However, RPCA is a transductive method and does not handle well the new samples which are not involved in the training procedure. Given a new datum, RPCA essentially needs to recalculate over all the data, resulting in high computational cost. So, RPCA is inappropriate for the applications that require fast online computation. To overcome this limitation, in this paper we propose an Inductive Robust Principal Component Analysis (IRPCA) method. Given a set of training data, unlike RPCA that targets on recovering the original data matrix, IRPCA aims at learning the underlying projection matrix, which can be used to efficiently remove the possible corruptions in any datum. The learning is done by solving a nuclear norm regularized minimization problem, which is convex and can be solved in polynomial time. Extensive experiments on a benchmark human face dataset and two video surveillance datasets show that IRPCA can not only be robust to gross corruptions, but also handle well the new data in an efficient way. Bing-Kun Bao, Guangcan Liu, Changsheng Xu, Shuicheng Yan |
IEEE Trans. Image Process. | 2 |
| 2012 | Saliency Detection by Multitask Sparsity PursuitabstractThis paper addresses the problem of detecting salient areas within natural images. We shall mainly study the problem under unsupervised setting, i.e., saliency detection without learning from labeled images. A solution of multitask sparsity pursuit is proposed to integrate multiple types of features for detecting saliency collaboratively. Given an image described by multiple features, its saliency map is inferred by seeking the consistently sparse elements from the joint decompositions of multiple-feature matrices into pairs of low-rank and sparse matrices. The inference process is formulated as a constrained nuclear norm and as an l(2, 1)-norm minimization problem, which is convex and can be solved efficiently with an augmented Lagrange multiplier method. Compared with previous methods, which usually make use of multiple features by combining the saliency maps obtained from individual features, the proposed method seamlessly integrates multiple features to produce jointly the saliency map with a single inference step and thus produces more accurate and reliable results. In addition to the unsupervised setting, the proposed method can be also generalized to incorporate the top-down priors obtained from supervised environment. Extensive experiments well validate its superiority over other state-of-the-art methods. Congyan Lang, Guangcan Liu, Jian Yu 0001, Shuicheng Yan |
IEEE Trans. Image Process. | 2 |
| 2011 | Multi-task low-rank affinity pursuit for image segmentationabstractThis paper investigates how to boost region-based image segmentation by pursuing a new solution to fuse multiple types of image features. A collaborative image segmentation framework, called multi-task low-rank affinity pursuit, is presented for such a purpose. Given an image described with multiple types of features, we aim at inferring a unified affinity matrix that implicitly encodes the segmentation of the image. This is achieved by seeking the sparsity-consistent low-rank affinities from the joint decompositions of multiple feature matrices into pairs of sparse and low-rank matrices, the latter of which is expressed as the production of the image feature matrix and its corresponding image affinity matrix. The inference process is formulated as a constrained nuclear norm and ℓ2;1-norm minimization problem, which is convex and can be solved efficiently with the Augmented Lagrange Multiplier method. Compared to previous methods, which are usually based on a single type of features, the proposed method seamlessly integrates multiple types of features to jointly produce the affinity matrix within a single inference step, and produces more accurate and reliable segmentation results. Experiments on the MSRC dataset and Berkeley segmentation dataset well validate the superiority of using multiple features over single feature and also the superiority of our method over conventional methods for feature fusion. Moreover, our method is shown to be very competitive while comparing to other state-of-the-art methods. Bin Cheng 0001, Guangcan Liu, Jingdong Wang 0001, ZhongYang Huang, Shuicheng Yan |
ICCV | 2 |
| 2011 | Latent Low-Rank Representation for subspace segmentation and feature extractionabstractLow-Rank Representation (LRR) [16, 17] is an effective method for exploring the multiple subspace structures of data. Usually, the observed data matrix itself is chosen as the dictionary, which is a key aspect of LRR. However, such a strategy may depress the performance, especially when the observations are insufficient and/or grossly corrupted. In this paper we therefore propose to construct the dictionary by using both observed and unobserved, hidden data. We show that the effects of the hidden data can be approximately recovered by solving a nuclear norm minimization problem, which is convex and can be solved efficiently. The formulation of the proposed method, called Latent Low-Rank Representation (LatLRR), seamlessly integrates subspace segmentation and feature extraction into a unified framework, and thus provides us with a solution for both subspace segmentation and feature extraction. As a subspace segmentation algorithm, LatLRR is an enhanced version of LRR and outperforms the state-of-the-art algorithms. Being an unsupervised feature extraction algorithm, LatLRR is able to robustly extract salient features from corrupted data, and thus can work much better than the benchmark that utilizes the original data vectors as features for classification. Compared to dimension reduction based methods, LatLRR is more robust to noise. Guangcan Liu, Shuicheng Yan |
ICCV | 1 |
| 2010 | Robust Subspace Segmentation by Low-Rank Representation
Guangcan Liu, Zhouchen Lin, Yong Yu 0001 |
ICML | 1 |
| 2010 | Unsupervised Object Segmentation with a Hybrid Graph Model (HGM)abstractIn this work, we address the problem of performing class-specific unsupervised object segmentation, i.e., automatic segmentation without annotated training images. Object segmentation can be regarded as a special data clustering problem where both class-specific information and local texture/color similarities have to be considered. To this end, we propose a hybrid graph model (HGM) that can make effective use of both symmetric and asymmetric relationship among samples. The vertices of a hybrid graph represent the samples and are connected by directed edges and/or undirected ones, which represent the asymmetric and/or symmetric relationship between them, respectively. When applied to object segmentation, vertices are superpixels, the asymmetric relationship is the conditional dependence of occurrence, and the symmetric relationship is the color/texture similarity. By combining the Markov chain formed by the directed subgraph and the minimal cut of the undirected subgraph, the object boundaries can be determined for each image. Using the HGM, we can conveniently achieve simultaneous segmentation and recognition by integrating both top-down and bottom-up information into a unified process. Experiments on 42 object classes (9,415 images in total) show promising results. Guangcan Liu, Zhouchen Lin, Yong Yu 0001, Xiaoou Tang |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2009 | Multi-output regression on the output manifold
Guangcan Liu, Zhouchen Lin, Yong Yu 0001 |
Pattern Recognit. | 1 |
| 2009 | Radon Representation-Based Feature Descriptor for Texture ClassificationabstractIn this paper, we aim to handle the intraclass variation resulting from the geometric transformation and the illumination change for more robust texture classification. To this end, we propose a novel feature descriptor called Radon representation-based feature descriptor (RRFD). RRFD converts the original pixel represented images into Radon-pixel images by using the Radon transform. The new Radon-pixel representation is more informative in geometry and has a much lower dimension. Subsequently, RRFD efficiently achieves affine invariance by projecting an image (or an image patch) from the space of Radon-pixel pairs onto an invariant feature space by using a ratiogram, i.e., the histogram of ratios between the areas of triangle pairs. The illumination invariance is also achieved by defining an illumination invariant distance metric on the invariant feature space. Comparing to the existing Radon transform-based texture features, which only achieve rotation and/or scaling invariance, RRFD achieves affine invariance. The experimental results on CUReT show that RRFD is a powerful feature descriptor that is suitable for texture classification. Guangcan Liu, Zhouchen Lin, Yong Yu 0001 |
IEEE Trans. Image Process. | 1 |
| 2005 | A Learning-Based Term-Weighting Approach for Information Retrieval
Guangcan Liu, Yong Yu 0001 |
AAAI | 1 |