VLDB 2026 Research / reviewers in the wild / expert
Li Liu 0004
dblp:33/4528-4
· DBLP profile ↗
144ranked-venue papers
22as first author
24since 2021 · last 2026
0009-0008-0974-5240ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 94 · 15 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 85 · 13 first-author · 8 since 2021Databases, data management, data science and information retrieval · 13 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2Systems, architecture and hardware · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Strip R-CNN: Large Strip Convolution for Remote Sensing Object DetectionabstractIn this paper, we show that current approaches using large square kernels or transformer-based global modeling aggregate contextual information uniformly across spatial dimensions, leading to feature dilution and localization errors for elongated targets. To mitigate this issue, we propose Strip R-CNN, the first work to systematically explore large strip convolutions for remote sensing object detection. Our key insight is that strip convolutions enable directional feature aggregation along the dominant spatial dimension of slender objects, reducing background interference while preserving essential geometric information. We design two core components: (i) StripNet, a backbone network employing sequential orthogonal large strip convolutions to capture anisotropic spatial patterns, and (ii) Strip Head, which enhances localization precision by incorporating strip convolutions into the detection head. Unlike previous large-kernel approaches that suffer from computational redundancy and isotropic limitations, our method achieves superior performance with remarkable efficiency. Extensive experiments on multiple benchmarks (DOTA, FAIR1M, HRSC2016, and DIOR) demonstrate significant improvements, with our 30M parameter model achieving 82.75% mAP on DOTA-v1.0, establishing a new state-of-the-art record while providing new insights into anisotropic feature learning for remote sensing applications. Xinbin Yuan, Zhaohui Zheng 0003, Yuxuan Li 0004, Xialei Liu, Li Liu 0004, Xiang Li 0041, Qibin Hou, Ming-Ming Cheng |
AAAI | 5 |
| 2025 | LSKNet: A Foundation Lightweight Backbone for Remote Sensing
Yuxuan Li 0004, Xiang Li 0041, Yimian Dai, Qibin Hou, Li Liu 0004, Yongxiang Liu, Ming-Ming Cheng, Jian Yang 0003 |
Int. J. Comput. Vis. | 5 |
| 2024 | SARDet-100K: Towards Open-Source Benchmark and ToolKit for Large-Scale SAR Object DetectionabstractSynthetic Aperture Radar (SAR) object detection has gained significant attention recently due to its irreplaceable all-weather imaging capabilities. However, this research field suffers from both limited public datasets (mostly comprising <2K images with only mono-category objects) and inaccessible source code. To tackle these challenges, we establish a new benchmark dataset and an open-source method for large-scale SAR object detection. Our dataset, SARDet-100K, is a result of intense surveying, collecting, and standardizing 10 existing SAR detection datasets, providing a large-scale and diverse dataset for research purposes. To the best of our knowledge, SARDet-100K is the first COCO-level large-scale multi-class SAR object detection dataset ever created. With this high-quality dataset, we conducted comprehensive experiments and uncovered a crucial challenge in SAR object detection: the substantial disparities between the pretraining on RGB datasets and finetuning on SAR datasets in terms of both data domain and model structure. To bridge these gaps, we propose a novel Multi-Stage with Filter Augmentation (MSFA) pretraining framework that tackles the problems from the perspective of data input, domain transition, and model migration. The proposed MSFA method significantly enhances the performance of SAR object detection models while demonstrating exceptional generalizability and flexibility across diverse models. This work aims to pave the way for further advancements in SAR object detection. The dataset and code is available at \url{https://github.com/zcablii/SARDet_100K}. Yuxuan Li 0004, Xiang Li 0041, Qibin Hou, Li Liu 0004, Ming-Ming Cheng, Jian Yang 0003 |
NeurIPS | 5 |
| 2023 | Data driven recurrent generative adversarial network for generalized zero shot image classification
Jie Zhang 0005, Shengbin Liao, Haofeng Zhang 0001, Yang Long 0001, Zheng Zhang 0006, Li Liu 0004 |
Inf. Sci. | 6 |
| 2023 | Normalization Techniques in Training DNNs: Methodology, Analysis and ApplicationabstractNormalization techniques are essential for accelerating the training and improving the generalization of deep neural networks (DNNs), and have successfully been used in various applications. This paper reviews and comments on the past, present and future of normalization methods in the context of DNN training. We provide a unified picture of the main motivation behind different approaches from the perspective of optimization, and present a taxonomy for understanding the similarities and differences between them. Specifically, we decompose the pipeline of the most representative normalizing activation methods into three components: the normalization area partitioning, normalization operation and normalization representation recovery. In doing so, we provide insight for designing new normalization technique. Finally, we discuss the current progress in understanding normalization methods, and provide a comprehensive review of the applications of normalization for particular tasks, in which it can effectively solve the key issues. Lei Huang 0015, Jie Qin 0004, Yi Zhou 0007, Fan Zhu 0001, Li Liu 0004, Ling Shao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2023 | Learning Multi-Attention Context Graph for Group-Based Re-IdentificationabstractLearning to re-identify or retrieve a group of people across non-overlapped camera systems has important applications in video surveillance. However, most existing methods focus on (single) person re-identification (re-id), ignoring the fact that people often walk in groups in real scenarios. In this work, we take a step further and consider employing context information for identifying groups of people, i.e., group re-id. On the one hand, group re-id is more challenging than single person re-id, since it requires both a robust modeling of local individual person appearance (with different illumination conditions, pose/viewpoint variations, and occlusions), as well as full awareness of global group structures (with group layout and group member variations). On the other hand, we believe that person re-id can be greatly enhanced by incorporating additional visual context from neighboring group members, a task which we formulate as group-aware (single) person re-id. In this paper, we propose a novel unified framework based on graph neural networks to simultaneously address the above two group-based re-id tasks, i.e., group re-id and group-aware person re-id. Specifically, we construct a context graph with group members as its nodes to exploit dependencies among different people. A multi-level attention mechanism is developed to formulate both intra-group and inter-group context, with an additional self-attention module for robust graph-level representations by attentively aggregating node-level features. The proposed model can be directly generalized to tackle group-aware person re-id using node-level representations. Meanwhile, to facilitate the deployment of deep learning models on these tasks, we build a new group re-id dataset which contains more than 3.8K images with 1.5K annotated groups, an order of magnitude larger than existing group re-id datasets. Extensive experiments on the novel dataset as well as three existing datasets clearly demonstrate the effectiveness of the proposed framework for both group-based re-id tasks. Yichao Yan, Jie Qin 0004, Bingbing Ni, Jiaxin Chen 0002, Li Liu 0004, Fan Zhu 0001, Wei-Shi Zheng 0001, Xiaokang Yang 0001, Ling Shao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2023 | Visual-Semantic Aligned Bidirectional Network for Zero-Shot LearningabstractZero-shot learning (ZSL) aims to recognize unknown categories that are unavailable during training. Recently, generative models have shown the potential to address this challenging problem by synthesizing unseen features conditioned on semantic embeddings such as attributes. However, unidirectional generative models cannot guarantee the effective coupling between visual and semantic spaces. To this end, we propose a visual-semantic aligned bidirectional network with cycle consistency to alleviate the gap between these two spaces, generating unseen features of high quality. More importantly, we incorporate two carefully designed strategies into our bidirectional framework to improve the overall ZSL performance. Specifically, we enhance the intra-domain class divergence in both visual and semantic spaces, and in the meantime, mitigate the inter-domain shift to preserve seen-unseen domain discrimination. Experimental results on four standard benchmarks show the superiority of our framework over existing state-of-the-art methods under both conventional and generalized ZSL settings. Xingsong Hou, Jie Qin 0004, Yuming Shen, Yang Long 0001, Li Liu 0004, Zhao Zhang 0001, Ling Shao 0001 |
IEEE Trans. Multim. | 6 |
| 2022 | Restore Globally, Refine Locally: A Mask-Guided Scheme to Accelerate Super-Resolution Networks
Xiaotao Hu, Jun Xu 0019, Shuhang Gu, Ming-Ming Cheng, Li Liu 0004 |
ECCV (19) | 5 |
| 2022 | Semi-supervised cross-modal hashing with multi-view graph representation
Haofeng Zhang 0001, Lunbo Li, Wankou Yang, Li Liu 0004 |
Inf. Sci. | 5 |
| 2022 | Learning discriminative and representative feature with cascade GAN for generalized zero-shot learning
Jingren Liu, Liyong Fu, Haofeng Zhang 0001, Qiaolin Ye, Wankou Yang, Li Liu 0004 |
Knowl. Based Syst. | 6 |
| 2022 | A Generalized Method for Binary Optimization: Convergence Analysis and ApplicationsabstractBinary optimization problems (BOPs) arise naturally in many fields, such as information retrieval, computer vision, and machine learning. Most existing binary optimization methods either use continuous relaxation which can cause large quantization errors, or incorporate a highly specific algorithm that can only be used for particular loss functions. To overcome these difficulties, we propose a novel generalized optimization method, named Alternating Binary Matrix Optimization (ABMO), for solving BOPs. ABMO can handle BOPs with/without orthogonality or linear constraints for a large class of loss functions. ABMO involves rewriting the binary, orthogonality and linear constraints for BOPs as an intersection of two closed sets, then iteratively dividing the original problems into several small optimization problems that can be solved as closed forms. To provide a strict theoretical convergence analysis, we add a sufficiently small perturbation and translate the original problem to an approximated problem whose feasible set is continuous. We not only provide rigorous mathematical proof for the convergence to a stationary and feasible point, but also derive the convergence rate of the proposed algorithm. The promising results obtained from four binary optimization tasks validate the superiority and the generality of ABMO compared with the state-of-the-art methods. Huan Xiong, Mengyang Yu, Li Liu 0004, Fan Zhu 0001, Jie Qin 0004, Fumin Shen, Ling Shao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Zero-shot learning via a specific rank-controlled semantic autoencoder
Yang Liu 0069, Xinbo Gao 0001, Jungong Han, Li Liu 0004, Ling Shao 0001 |
Pattern Recognit. | 4 |
| 2022 | Generalized Zero-Shot Learning With Multiple Graph Adaptive Generative NetworksabstractGenerative adversarial networks (GANs) for (generalized) zero-shot learning (ZSL) aim to generate unseen image features when conditioned on unseen class embeddings, each of which corresponds to one unique category. Most existing works on GANs for ZSL generate features by merely feeding the seen image feature/class embedding (combined with random Gaussian noise) pairs into the generator/discriminator for a two-player minimax game. However, the structure consistency of the distributions among the real/fake image features, which may shift the generated features away from their real distribution to some extent, is seldom considered. In this paper, to align the weights of the generator for better structure consistency between real/fake features, we propose a novel multigraph adaptive GAN (MGA-GAN). Specifically, a Wasserstein GAN equipped with a classification loss is trained to generate discriminative features with structure consistency. MGA-GAN leverages the multigraph similarity structures between sliced seen real/fake feature samples to assist in updating the generator weights in the local feature manifold. Moreover, correlation graphs for the whole real/fake features are adopted to guarantee structure correlation in the global feature manifold. Extensive evaluations on four benchmarks demonstrate well the superiority of MGA-GAN over its state-of-the-art counterparts. Guosen Xie, Zheng Zhang 0006, Guoshuai Liu, Fan Zhu 0001, Li Liu 0004, Ling Shao 0001, Xuelong Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2021 | Group Whitening: Balancing Learning Efficiency and Representational CapacityabstractBatch normalization (BN) is an important technique commonly incorporated into deep learning models to perform standardization within mini-batches. The merits of BN in improving a model’s learning efficiency can be further amplified by applying whitening, while its drawbacks in estimating population statistics for inference can be avoided through group normalization (GN). This paper proposes group whitening (GW), which exploits the advantages of the whitening operation and avoids the disadvantages of normalization within mini-batches. In addition, we analyze the constraints imposed on features by normalization, and show how the batch size (group number) affects the performance of batch (group) normalized networks, from the perspective of model’s representational capacity. This analysis provides theoretical guidance for applying GW in practice. Finally, we apply the proposed GW to ResNet and ResNeXt architectures and conduct experiments on the ImageNet and COCO benchmarks. Results show that GW consistently improves the performance of different architectures, with absolute gains of 1.02% ∼ 1.49% in top-1 accuracy on ImageNet and 1.82% ∼ 3.21% in bounding box AP on COCO. Lei Huang 0015, Yi Zhou 0007, Li Liu 0004, Fan Zhu 0001, Ling Shao 0001 |
CVPR | 3 |
| 2021 | Anchor-Free Person SearchabstractPerson search aims to simultaneously localize and identify a query person from realistic, uncropped images, which can be regarded as the unified task of pedestrian detection and person re-identification (re-id). Most existing works employ two-stage detectors like Faster-RCNN, yielding encouraging accuracy but with high computational overhead. In this work, we present the Feature-Aligned Person Search Network (AlignPS), the first anchor-free framework to efficiently tackle this challenging task. AlignPS explicitly addresses the major challenges, which we summarize as the misalignment issues in different levels (i.e., scale, region, and task), when accommodating an anchor-free detector for this task. More specifically, we propose an aligned feature aggregation module to generate more discriminative and robust feature embeddings by following a "re-id first" principle. Such a simple design directly improves the baseline anchor-free model on CUHK-SYSU by more than 20% in mAP. Moreover, AlignPS outperforms state-of-the-art two-stage methods, with a higher speed. The code is available at https://github.com/daodaofr/AlignPS. Yichao Yan, Jinpeng Li 0004, Jie Qin 0004, Song Bai 0001, Shengcai Liao, Li Liu 0004, Fan Zhu 0001, Ling Shao 0001 |
CVPR | 6 |
| 2021 | Near-Real Feature Generative Network for Generalized Zero-Shot LearningabstractDue to the powerful feature synthesis ability, Generative Adversarial Networks (GAN) is well adapted to the Generalized Zero-Shot Learning (GZSL) task and has achieved great success. Most GAN models for GZSL usually employ random noise with normal distribution to synthesize unseen samples. However, the generated samples often have the same normal distribution as the input noise, which is unrealistic in most circumstances. Therefore, in this paper, we consider that the distribution of unseen classes should be follow that of seen classes and propose a near-real feature generative network (NereNet), which utilizes the most semantically similar seen samples to generate the noise for the unseen classes. Specifically, we first calculate the most similar seen classes for the unseen classes, and then train an encoder network to generate the corresponding noise, which is subsequently combined with the unseen classes attributes to generate unseen samples with GAN. Extensive experiments are conducted on four datasets, and the results demonstrate the effectiveness of our proposed method. Jingren Liu, Haoyue Bai 0003, Haofeng Zhang 0001, Li Liu 0004 |
ICME | 4 |
| 2021 | Attention-Guided Semantic Hashing for Unsupervised Cross-Modal RetrievalabstractRecently, due to the low storage consumption and high search efficiency of hashing methods and the powerful feature extraction capability of deep neural networks, deep cross-modal hashing has received extensive attention in the field of multi-media retrieval. However, existing methods tend to ignore the latent relationships between heterogeneous data when learning a common semantic subspace, and cannot retain more important semantic information when mining deep correlations. In this paper, an attention mechanism which focuses on the characteristics of the associated features is employed to propose an attention-aware semantic fusion matrix that integrates important information from different modalities. We introduce a novel network that can pass the extracted features through the attention module to efficiently encode rich and relevant features, and can also generate hash codes under the self-supervision of the proposed attention-aware semantic fusion matrix. Our experimental results and detailed analysis prove that our method can achieve better retrieval performance on the three popular datasets, compared with the recent unsupervised cross-modal hashing methods. Haofeng Zhang 0001, Lunbo Li, Li Liu 0004 |
ICME | 4 |
| 2021 | Clustering-driven Deep Adversarial Hashing for scalable unsupervised cross-modal retrieval
Haofeng Zhang 0001, Lunbo Li, Zheng Zhang 0006, Debao Chen, Li Liu 0004 |
Neurocomputing | 6 |
| 2021 | Sparse graph based self-supervised hashing for scalable image retrieval
Haofeng Zhang 0001, Zheng Zhang 0006, Li Liu 0004, Ling Shao 0001 |
Inf. Sci. | 4 |
| 2021 | Learning Deformable and Attentive Network for image restoration
Xingsong Hou, Yujie Dun, Jie Qin 0004, Li Liu 0004, Xueming Qian, Ling Shao 0001 |
Knowl. Based Syst. | 5 |
| 2021 | A plug-in attribute correction module for generalized zero-shot learning
Haofeng Zhang 0001, Haoyue Bai 0003, Yang Long 0001, Li Liu 0004, Ling Shao 0001 |
Pattern Recognit. | 4 |
| 2021 | Deep Unsupervised Self-Evolutionary Hashing for Image RetrievalabstractHashing methods have proven to be effective in the field of large-scale image retrieval. In recent years, the performance of hashing algorithms based on deep learning has greatly exceeded that of non-deep methods. However, most of the outstanding hashing methods are supervised models that heavily rely on annotated labels. In order to circumvent the huge overhead of labeling large-scale datasets, some unsupervised hashing algorithms have been proposed, such as pseudo labels and pseudo pairs. Since the image labels are strictly unavailable, some hyper-parameters in these methods are difficult to be selected, e.g., the final result is very sensitive to the picked number of categories or the chosen threshold of similarity for pairs. In addition, the calculation of pseudo-labels in high-dimensional space is not only computationally complex, but also has low precision. Therefore, in order to alleviate these issues in this paper, we propose a simple but effective Deep Unsupervised Self-evolutionary Hashing (DUSH) algorithm, which utilizes a curriculum learning strategy to iteratively select pseudo pairs from easy to hard in low dimensional Hamming space. Extensive experiments are conducted on four popular datasets, including two single-label datasets and two multi-label datasets, and the results show that our method can significantly outperform the state-of-the-art methods. Haofeng Zhang 0001, Yazhou Yao, Zheng Zhang 0006, Li Liu 0004, Jian Zhang 0002, Ling Shao 0001 |
IEEE Trans. Multim. | 5 |
| 2021 | Internal and external memory set containment join
Chengcheng Yang, Dong Deng 0001, Shuo Shang, Fan Zhu 0001, Li Liu 0004, Ling Shao 0001 |
VLDB J. | 5 |
| 2021 | Correction to: Internal and external memory set containment join
Chengcheng Yang, Dong Deng 0001, Shuo Shang, Fan Zhu 0001, Li Liu 0004, Ling Shao 0001 |
VLDB J. | 5 |
| 2020 | Controllable Orthogonalization in Training DNNsabstractOrthogonality is widely used for training deep neural networks (DNNs) due to its ability to maintain all singular values of the Jacobian close to 1 and reduce redundancy in representation. This paper proposes a computationally efficient and numerically stable orthogonalization method using Newton's iteration (ONI), to learn a layer-wise orthogonal weight matrix in DNNs. ONI works by iteratively stretching the singular values of a weight matrix towards 1. This property enables it to control the orthogonality of a weight matrix by its number of iterations. We show that our method improves the performance of image classification networks by effectively controlling the orthogonality to provide an optimal tradeoff between optimization benefits and representational capacity reduction. We also show that ONI stabilizes the training of generative adversarial networks (GANs) by maintaining the Lipschitz continuity of a network, similar to spectral normalization (SN), and further outperforms SN by providing controllable orthogonality. Lei Huang 0015, Li Liu 0004, Fan Zhu 0001, Diwen Wan, Zehuan Yuan, Bo Li 0026, Ling Shao 0001 |
CVPR | 2 |
| 2020 | An Investigation Into the Stochasticity of Batch WhiteningabstractBatch Normalization (BN) is extensively employed in various network architectures by performing standardization within mini-batches. A full understanding of the process has been a central target in the deep learning communities. Unlike existing works, which usually only analyze the standardization operation, this paper investigates the more general Batch Whitening (BW). Our work originates from the observation that while various whitening transformations equivalently improve the conditioning, they show significantly different behaviors in discriminative scenarios and training Generative Adversarial Networks (GANs). We attribute this phenomenon to the stochasticity that BW introduces. We quantitatively investigate the stochasticity of different whitening transformations and show that it correlates well with the optimization behaviors during training. We also investigate how stochasticity relates to the estimation of population statistics during inference. Based on our analysis, we provide a framework for designing and comparing BW algorithms in different scenarios. Our proposed BW algorithm improves the residual networks by a significant margin on ImageNet classification. Besides, we show that the stochasticity of BW can improve the GAN's performance with, however, the sacrifice of the training stability. Lei Huang 0015, Yi Zhou 0007, Fan Zhu 0001, Li Liu 0004, Ling Shao 0001 |
CVPR | 5 |
| 2020 | Auto-Encoding Twin-Bottleneck HashingabstractConventional unsupervised hashing methods usually take advantage of similarity graphs, which are either pre-computed in the high-dimensional space or obtained from random anchor points. On the one hand, existing methods uncouple the procedures of hash function learning and graph construction. On the other hand, graphs empirically built upon original data could introduce biased prior knowledge of data relevance, leading to sub-optimal retrieval performance. In this paper, we tackle the above problems by proposing an efficient and adaptive code-driven graph, which is updated by decoding in the context of an auto-encoder. Specifically, we introduce into our framework twin bottlenecks (i.e., latent variables) that exchange crucial information collaboratively. One bottleneck (i.e., binary codes) conveys the high-level intrinsic data structure captured by the code-driven graph to the other (i.e., continuous variables for low-level detail information), which in turn propagates the updated network feedback for the encoder to learn more discriminative binary codes. The auto-encoding learning objective literally rewards the code-driven graph to learn an optimal encoder. Moreover, the proposed model can be simply optimized by gradient descent without violating the binary constraints. Experiments on benchmarked datasets clearly show the superiority of our framework over the state-of-the-art hashing methods. Our source code can be found at https://github.com/ymcidence/TBH. Yuming Shen, Jie Qin 0004, Jiaxin Chen 0002, Mengyang Yu, Li Liu 0004, Fan Zhu 0001, Fumin Shen, Ling Shao 0001 |
CVPR | 5 |
| 2020 | Learning Multi-Granular Hypergraphs for Video-Based Person Re-IdentificationabstractVideo-based person re-identification (re-ID) is an important research topic in computer vision. The key to tackling the challenging task is to exploit both spatial and temporal clues in video sequences. In this work, we propose a novel graph-based framework, namely Multi-Granular Hypergraph (MGH), to pursue better representational capabilities by modeling spatiotemporal dependencies in terms of multiple granularities. Specifically, hypergraphs with different spatial granularities are constructed using various levels of part-based features across the video sequence. In each hypergraph, different temporal granularities are captured by hyperedges that connect a set of graph nodes (i.e., part-based features) across different temporal ranges. Two critical issues (misalignment and occlusion) are explicitly addressed by the proposed hypergraph propagation and feature aggregation schemes. Finally, we further enhance the overall video representation by learning more diversified graph-level representations of multiple granularities based on mutual information minimization. Extensive experiments on three widely-adopted benchmarks clearly demonstrate the effectiveness of the proposed framework. Notably, 90.0% top-1 accuracy on MARS is achieved using MGH, outperforming the state-of-the-arts. Yichao Yan, Jie Qin 0004, Jiaxin Chen 0002, Li Liu 0004, Fan Zhu 0001, Ying Tai, Ling Shao 0001 |
CVPR | 4 |
| 2020 | Learning Attentive and Hierarchical Representations for 3D Shape Recognition
Jiaxin Chen 0002, Jie Qin 0004, Yuming Shen, Li Liu 0004, Fan Zhu 0001, Ling Shao 0001 |
ECCV (15) | 4 |
| 2020 | Layer-Wise Conditioning Analysis in Exploring the Learning Dynamics of DNNs
Lei Huang 0015, Jie Qin 0004, Li Liu 0004, Fan Zhu 0001, Ling Shao 0001 |
ECCV (2) | 3 |
| 2020 | Invertible Zero-Shot Recognition Flows
Yuming Shen, Jie Qin 0004, Lei Huang 0015, Li Liu 0004, Fan Zhu 0001, Ling Shao 0001 |
ECCV (16) | 4 |
| 2020 | Region Graph Embedding Network for Zero-Shot Learning
Guosen Xie, Li Liu 0004, Fan Zhu 0001, Fang Zhao 0006, Zheng Zhang 0006, Yazhou Yao, Jie Qin 0004, Ling Shao 0001 |
ECCV (4) | 2 |
| 2020 | On the Number of Linear Regions of Convolutional Neural NetworksabstractOne fundamental problem in deep learning is understanding the outstanding performance of deep Neural Networks (NNs) in practice. One explanation for the superiority of NNs is that they can realize a large class of complicated functions, i.e., they have powerful expressivity. The expressivity of a ReLU NN can be quantified by the maximal number of linear regions it can separate its input space into. In this paper, we provide several mathematical results needed for studying the linear regions of CNNs, and use them to derive the maximal and average numbers of linear regions for one-layer ReLU CNNs. Furthermore, we obtain upper and lower bounds for the number of linear regions of multi-layer ReLU CNNs. Our results suggest that deeper CNNs have more powerful expressivity than their shallow counterparts, while CNNs have more expressivity than fully-connected NNs per parameter. Huan Xiong, Lei Huang 0015, Mengyang Yu, Li Liu 0004, Fan Zhu 0001, Ling Shao 0001 |
ICML | 4 |
| 2020 | Improved Residual Networks for Image and Video RecognitionabstractResidual networks (ResNets) represent a powerful type of convolutional neural network (CNN) architecture, widely adopted and used in various tasks. In this work we propose an improved version of ResNets. Our proposed improvements address all three main components of a ResNet: the flow of information through the network layers, the residual building block, and the projection shortcut. We are able to show consistent improvements in accuracy and learning convergence over the baseline. For instance, on ImageNet dataset, using the ResNet with 50 layers, for top-1 accuracy we can report a 1.19% improvement over the baseline in one setting and around 2% boost in another. Importantly, these improvements are obtained without increasing the model complexity. Our proposed approach allows us to train extremely deep networks, while the baseline shows severe optimization issues. We report results on three tasks over six datasets: image classification (ImageNet, CIFAR-10 and CIFAR-100), object detection (COCO) and video action recognition (Kinetics-400 and Something-Something-v2). In the deep learning era, we establish a new milestone for the depth of a CNN. We successfully train a 404-layer deep CNN on the ImageNet dataset and a 3002-layer network on CIFAR-10 and CIFAR-100, while the baseline is not able to converge at such extreme depths. Code and models are publicly available at: https://github.com/iduta/iresnet. I. C. Duta, Li Liu 0004, Fan Zhu 0001, Ling Shao 0001 |
ICPR | 2 |
| 2020 | Set and Rebase: Determining the Semantic Graph Connectivity for Unsupervised Cross-Modal HashingabstractThe label-free nature of unsupervised cross-modal hashing hinders models from exploiting the exact semantic data similarity. Existing research typically simulates the semantics by a heuristic geometric prior in the original feature space. However, this introduces heavy bias into the model as the original features are not fully representing the underlying multi-view data relations. To address the problem above, in this paper, we propose a novel unsupervised hashing method called Semantic-Rebased Cross-modal Hashing (SRCH). A novel ‘Set-and-Rebase’ process is defined to initialize and update the cross-modal similarity graph of training data. In particular, we set the graph according to the intra-modal feature geometric basis and then alternately rebase it to update the edges within according to the hashing results. We develop an alternating optimization routine to rebase the graph and train the hashing auto-encoders with closed-form solutions so that the overall framework is efficiently trained. Our experimental results on benchmarked datasets demonstrate the superiority of our model against state-of-the-art algorithms. Yuming Shen, Haofeng Zhang 0001, Yazhou Yao, Li Liu 0004 |
IJCAI | 5 |
| 2020 | Deep Local Binary Coding for Person Re-Identification by Delving into the DetailsabstractPerson re-identification (ReID) has recently received extensive research interests due to its diverse applications in multimedia analysis and computer vision. However, the majority of existing works focus on improving matching accuracy, while ignoring matching efficiency. In this work, we present a novel binary representation learning framework for efficient person ReID, namely Deep Local Binary Coding (DLBC). Different from existing deep binary ReID approaches, DLBC attempts to learn discriminative binary codes by explicitly interacting with local visual details. Specifically, DLBC first extracts a set of local features from spatially salient regions of pedestrian images. Subsequently, DLBC formulates a new binary-local semantic mutual information (BSMI) maximization term, based on which a self-lifting (SL) block is built to further exploit the semantic importance of local features. The BSMI term together with the SL block simultaneously enhances the dependency of binary codes on selected local features as well as their robustness to cross-view visual inconsistency. In addition, an efficient optimizing method is developed to train the proposed deep models with orthogonal and binary constraints. Extensive experiments reveal that DLBC significantly minimizes the accuracy gap between binary ReID methods and the state-of-the-art real-valued ones, whilst remarkably reducing query time and memory cost. Jiaxin Chen 0002, Jie Qin 0004, Yichao Yan, Lei Huang 0015, Li Liu 0004, Fan Zhu 0001, Ling Shao 0001 |
ACM Multimedia | 5 |
| 2020 | Semantic-rebased cross-modal hashing for scalable unsupervised text-visual retrieval
Yuming Shen, Haofeng Zhang 0001, Li Liu 0004 |
Inf. Process. Manag. | 4 |
| 2020 | Projection based weight normalization: Efficient method for optimization on oblique manifold in DNNs
Lei Huang 0015, Xianglong Liu 0001, Jie Qin 0004, Fan Zhu 0001, Li Liu 0004, Ling Shao 0001 |
Pattern Recognit. | 5 |
| 2020 | Deep quantization generative networks
Diwen Wan, Fumin Shen, Li Liu 0004, Fan Zhu 0001, Lei Huang 0015, Mengyang Yu, Heng Tao Shen, Ling Shao 0001 |
Pattern Recognit. | 3 |
| 2020 | Deep transductive network for generalized zero shot learning
Haofeng Zhang 0001, Li Liu 0004, Yang Long 0001, Zheng Zhang 0006, Ling Shao 0001 |
Pattern Recognit. | 2 |
| 2020 | Graph Convolutional Network HashingabstractRecently, graph-based hashing that learns similarity-preserving binary codes via an affinity graph has been extensively studied for large-scale image retrieval. However, most graph-based hashing methods resort to intractable binary quadratic programs, making them unscalable to massive data. In this paper, we propose a novel graph convolutional network-based hashing framework, dubbed GCNH, which directly carries out spectral convolution operations on both an image set and an affinity graph built over the set, naturally yielding similarity-preserving binary embedding. GCNH fundamentally differs from conventional graph hashing methods which adopt an affinity graph as the only learning guidance in an objective function to pursue the binary embedding. As the core ingredient of GCNH, we introduce an intuitive asymmetric graph convolutional (AGC) layer to simultaneously convolve the anchor graph, input data, and convolutional filters. By virtue of the AGC layer, GCNH well addresses the issues of scalability and out-of-sample extension when leveraging affinity graphs for hashing. As a use case of our GCNH, we particularly study the semisupervised hashing scenario in this paper. Comprehensive image retrieval evaluations on the CIFAR-10, NUS-WIDE, and ImageNet datasets demonstrate the consistent advantages of GCNH over the state-of-the-art methods given limited labeled data. Xiang Zhou 0008, Fumin Shen, Li Liu 0004, Wei Liu 0005, Liqiang Nie, Yang Yang 0002, Heng Tao Shen |
IEEE Trans. Cybern. | 3 |
| 2020 | Zero-VAE-GAN: Generating Unseen Features for Generalized and Transductive Zero-Shot LearningabstractZero-shot learning (ZSL) is a challenging task due to the lack of unseen class data during training. Existing works attempt to establish a mapping between the visual and class spaces through a common intermediate semantic space. The main limitation of existing methods is the strong bias towards seen class, known as the domain shift problem, which leads to unsatisfactory performance in both conventional and generalized ZSL tasks. To tackle this challenge, we propose to convert ZSL to the conventional supervised learning by generating features for unseen classes. To this end, a joint generative model that couples variational autoencoder (VAE) and generative adversarial network (GAN), called Zero-VAE-GAN, is proposed to generate high-quality unseen features. To enhance the class-level discriminability, an adversarial categorization network is incorporated into the joint framework. Besides, we propose two self-training strategies to augment unlabeled unseen features for the transductive extension of our model, addressing the domain shift problem to a large extent. Experimental results on five standard benchmarks and a large-scale dataset demonstrate the superiority of our generative model over the state-of-the-art methods for conventional, especially generalized ZSL tasks. Moreover, the further improvement of the transductive setting demonstrates the effectiveness of the proposed self-training strategies. Xingsong Hou, Jie Qin 0004, Jiaxin Chen 0002, Li Liu 0004, Fan Zhu 0001, Zhao Zhang 0001, Ling Shao 0001 |
IEEE Trans. Image Process. | 5 |
| 2020 | NLH: A Blind Pixel-Level Non-Local Method for Real-World Image DenoisingabstractNon-local self similarity (NSS) is a powerful prior of natural images for image denoising. Most of existing denoising methods employ similar patches, which is a patch-level NSS prior. In this paper, we take one step forward by introducing a pixel-level NSS prior, i.e., searching similar pixels across a non-local region. This is motivated by the fact that finding closely similar pixels is more feasible than similar patches in natural images, which can be used to enhance image denoising performance. With the introduced pixel-level NSS prior, we propose an accurate noise level estimation method, and then develop a blind image denoising method based on the lifting Haar transform and Wiener filtering techniques. Experiments on benchmark datasets demonstrate that, the proposed method achieves much better performance than previous non-deep methods, and is still competitive with existing state-of-the-art deep learning based methods on real-world image denoising. The code is publicly available athttps://github.com/njusthyk1972/NLH. Yingkun Hou, Jun Xu 0019, Mingxia Liu 0001, Guanghai Liu 0001, Li Liu 0004, Fan Zhu 0001, Ling Shao 0001 |
IEEE Trans. Image Process. | 5 |
| 2020 | Noisy-as-Clean: Learning Self-Supervised Denoising From Corrupted ImageabstractSupervised deep networks have achieved promising performance on image denoising, by learning image priors and noise statistics on plenty pairs of noisy and clean images. Unsupervised denoising networks are trained with only noisy images. However, for an unseen corrupted image, both supervised and unsupervised networks ignore either its particular image prior, the noise statistics, or both. That is, the networks learned from external images inherently suffer from a domain gap problem: the image priors and noise statistics are very different between the training and test images. This problem becomes more clear when dealing with the signal dependent realistic noise. To circumvent this problem, in this work, we propose a novel "Noisy-As-Clean" (NAC) strategy of training self-supervised denoising networks. Specifically, the corrupted test image is directly taken as the "clean" target, while the inputs are synthetic images consisted of this corrupted image and a second yet similar corruption. A simple but useful observation on our NAC is: as long as the noise is weak, it is feasible to learn a self-supervised network only with the corrupted image, approximating the optimal parameters of a supervised network learned with pairs of noisy and clean images. Experiments on synthetic and realistic noise removal demonstrate that, the DnCNN and ResNet networks trained with our self-supervised NAC strategy achieve comparable or better performance than the original ones and previous supervised/unsupervised/self-supervised networks. The code is publicly available at https://github.com/csjunxu/Noisy-As-Clean. Jun Xu 0019, Ming-Ming Cheng, Li Liu 0004, Fan Zhu 0001, Zhou Xu 0003, Ling Shao 0001 |
IEEE Trans. Image Process. | 4 |
| 2020 | STAR: A Structure and Texture Aware Retinex ModelabstractRetinex theory is developed mainly to decompose an image into the illumination and reflectance components by analyzing local image derivatives. In this theory, larger derivatives are attributed to the changes in reflectance, while smaller derivatives are emerged in the smooth illumination. In this paper, we utilize exponentiated local derivatives (with an exponent γ) of an observed image to generate its structure map and texture map. The structure map is produced by been amplified with γ > 1, while the texture map is generated by been shrank with γ < 1. To this end, we design exponential filters for the local derivatives, and present their capability on extracting accurate structure and texture maps, influenced by the choices of exponents γ. The extracted structure and texture maps are employed to regularize the illumination and reflectance components in Retinex decomposition. A novel Structure and Texture Aware Retinex (STAR) model is further proposed for illumination and reflectance decomposition of a single image. We solve the STAR model by an alternating optimization algorithm. Each sub-problem is transformed into a vectorized least squares regression, with closed-form solutions. Comprehensive experiments on commonly tested datasets demonstrate that, the proposed STAR model produce better quantitative and qualitative performance than previous competing methods, on illumination and reflectance decomposition, low-light image enhancement, and color correction. The code is publicly available at https://github.com/csjunxu/STAR. Jun Xu 0019, Yingkun Hou, Dongwei Ren, Li Liu 0004, Fan Zhu 0001, Mengyang Yu, Haoqian Wang, Ling Shao 0001 |
IEEE Trans. Image Process. | 4 |
| 2020 | Towards Automatic Construction of Diverse, High-Quality Image DatasetsabstractThe availability of labeled image datasets has been shown critical for high-level image understanding, which continuously drives the progress of feature designing and models developing. However, constructing labeled image datasets is laborious and monotonous. To eliminate manual annotation, in this work, we propose a novel image dataset construction framework by employing multiple textual queries. We aim at collecting diverse and accurate images for given queries from the Web. Specifically, we formulate noisy textual queries removing and noisy images filtering as a multi-view and multi-instance learning problem separately. Our proposed approach not only improves the accuracy but also enhances the diversity of the selected images. To verify the effectiveness of our proposed approach, we construct an image dataset with 100 categories. The experiments show significant performance gains by using the generated data of our approach on several tasks, such as image classification, cross-dataset generalization, and object detection. The proposed method also consistently outperforms existing weakly supervised and web-supervised approaches. Yazhou Yao, Jian Zhang 0002, Fumin Shen, Li Liu 0004, Fan Zhu 0001, Dongxiang Zhang, Heng Tao Shen |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2020 | Approximate Kernel Selection via Matrix ApproximationabstractKernel selection is of fundamental importance for the generalization of kernel methods. This article proposes an approximate approach for kernel selection by exploiting the approximability of kernel selection and the computational virtue of kernel matrix approximation. We define approximate consistency to measure the approximability of the kernel selection problem. Based on the analysis of approximate consistency, we solve the theoretical problem of whether, under what conditions, and at what speed, the approximate criterion is close to the accurate one, establishing the foundations of approximate kernel selection. We introduce two selection criteria based on error estimation and prove the approximate consistency of the multilevel circulant matrix (MCM) approximation and Nyström approximation under these criteria. Under the theoretical guarantees of the approximate consistency, we design approximate algorithms for kernel selection, which exploits the computational advantages of the MCM and Nyström approximations to conduct kernel selection in a linear or quasi-linear complexity. We experimentally validate the theoretical results for the approximate consistency and evaluate the effectiveness of the proposed kernel selection algorithms. Lizhong Ding 0001, Shizhong Liao, Yong Liu 0018, Li Liu 0004, Fan Zhu 0001, Yazhou Yao, Ling Shao 0001, Xin Gao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2020 | SRSC: Selective, Robust, and Supervised Constrained Feature Representation for Image ClassificationabstractFeature representation learning, an emerging topic in recent years, has achieved great progress. Powerful learned features can lead to excellent classification accuracy. In this article, a selective and robust feature representation framework with a supervised constraint (SRSC) is presented. SRSC seeks a selective, robust, and discriminative subspace by transforming the original feature space into the category space. Particularly, we add a selective constraint to the transformation matrix (or classifier parameter) that can select discriminative dimensions of the input samples. Moreover, a supervised regularization is tailored to further enhance the discriminability of the subspace. To relax the hard zero-one label matrix in the category space, an additional error term is also incorporated into the framework, which can lead to a more robust transformation matrix. SRSC is formulated as a constrained least square learning (feature transforming) problem. For the SRSC problem, an inexact augmented Lagrange multiplier method (ALM) is utilized to solve it. Extensive experiments on several benchmark data sets adequately demonstrate the effectiveness and superiority of the proposed method. The proposed SRSC approach has achieved better performances than the compared counterpart methods. Guosen Xie, Zheng Zhang 0006, Li Liu 0004, Fan Zhu 0001, Xu-Yao Zhang, Ling Shao 0001, Xuelong Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2020 | Exploiting Web Images for Multi-Output Classification: From Category to SubcategoriesabstractStudies present that dividing categories into subcategories contributes to better image classification. Existing image subcategorization works relying on expert knowledge and labeled images are both time-consuming and labor-intensive. In this article, we propose to select and subsequently classify images into categories and subcategories. Specifically, we first obtain a list of candidate subcategory labels from untagged corpora. Then, we purify these subcategory labels through calculating the relevance to the target category. To suppress the search error and noisy subcategory label-induced outlier images, we formulate outlier images removing and the optimal classification models learning as a unified problem to jointly learn multiple classifiers, where the classifier for a category is obtained by combining multiple subcategory classifiers. Compared with the existing subcategorization works, our approach eliminates the dependence on expert knowledge and labeled images. Extensive experiments on image categorization and subcategorization demonstrate the superiority of our proposed approach. Yazhou Yao, Fumin Shen, Guosen Xie, Li Liu 0004, Fan Zhu 0001, Jian Zhang 0002, Heng Tao Shen |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2019 | Deep Sketch-Shape Hashing With Segmented 3D Stochastic ViewingabstractSketch-based 3D shape retrieval has been extensively studied in recent works, most of which focus on improving the retrieval accuracy, whilst neglecting the efficiency. In this paper, we propose a novel framework for efficient sketch-based 3D shape retrieval, i.e., Deep Sketch-Shape Hashing (DSSH), which tackles the challenging problem from two perspectives. Firstly, we propose an intuitive 3D shape representation method to deal with unaligned shapes with arbitrary poses. Specifically, the proposed Segmented Stochastic-viewing Shape Network models discriminative 3D representations by a set of 2D images rendered from multiple views, which are stochastically selected from non-overlapping spatial segments of a 3D sphere. Secondly, Batch-Hard Binary Coding (BHBC) is developed to learn semantics-preserving compact binary codes by mining the hardest samples. The overall framework is jointly learned by developing an alternating iteration algorithm. Extensive experimental results on three benchmarks show that DSSH improves both the retrieval efficiency and accuracy remarkably, compared to the state-of-the-art methods. Jiaxin Chen 0002, Jie Qin 0004, Li Liu 0004, Fan Zhu 0001, Fumin Shen, Jin Xie 0001, Ling Shao 0001 |
CVPR | 3 |
| 2019 | Iterative Normalization: Beyond Standardization Towards Efficient WhiteningabstractBatch Normalization (BN) is ubiquitously employed for accelerating neural network training and improving the generalization capability by performing standardization within mini-batches. Decorrelated Batch Normalization (DBN) further boosts the above effectiveness by whitening. However, DBN relies heavily on either a large batch size, or eigen-decomposition that suffers from poor efficiency on GPUs. We propose Iterative Normalization (IterNorm), which employs Newton’s iterations for much more efficient whitening, while simultaneously avoiding the eigen-decomposition. Furthermore, we develop a comprehensive study to show IterNorm has better trade-off between optimization and generalization, with theoretical and experimental support. To this end, we exclusively introduce Stochastic Normalization Disturbance (SND), which measures the inherent stochastic uncertainty of samples when applied to normalization operations. With the support of SND, we provide natural explanations to several phenomena from the perspective of optimization, e.g., why group-wise whitening of DBN generally outperforms full-whitening and why the accuracy of BN degenerates with reduced batch sizes. We demonstrate the consistently improved performance of IterNorm with extensive experiments on CIFAR-10 and ImageNet over BN and DBN. Lei Huang 0015, Yi Zhou 0007, Fan Zhu 0001, Li Liu 0004, Ling Shao 0001 |
CVPR | 4 |
| 2019 | Building Detail-Sensitive Semantic Segmentation Networks With Polynomial PoolingabstractSemantic segmentation is an important computer vision task, which aims to allocate a semantic label to each pixel in an image. When training a segmentation model, it is common to fine-tune a classification network pre-trained on a large-scale dataset. However, as an intrinsic property of the classification model, invariance to spatial perturbation resulting from the lose of detail-sensitivity prevents segmentation networks from achieving high performance. The use of standard poolings is one of the key factors for this invariance. The most common standard poolings are max and average pooling. Max pooling can increase both the invariance to spatial perturbations and the non-linearity of the networks. Average pooling, on the other hand, is sensitive to spatial perturbations, but is a linear function. For semantic segmentation, we prefer both the preservation of detailed cues within a local feature region and non-linearity that increases a network's functional complexity. In this work, we propose a polynomial pooling (P-pooling) function that finds an intermediate form between max and average pooling to provide an optimally balanced and self-adjusted pooling strategy for semantic segmentation. The P-pooling is differentiable and can be applied into a variety of pre-trained networks. Extensive studies on the PASCAL VOC, Cityscapes and ADE20k datasets demonstrate the superiority of P-pooling over other poolings. Experiments on various network architectures and state-of-the-art training strategies also show that models with P-pooling layers consistently outperform those directly fine-tuned using pre-trained classification models. Zhen Wei 0001, Jingyi Zhang 0005, Li Liu 0004, Fan Zhu 0001, Fumin Shen, Yi Zhou 0007, Si Liu 0001, Yao Sun 0004, Ling Shao 0001 |
CVPR | 3 |
| 2019 | Attentive Region Embedding Network for Zero-Shot LearningabstractZero-shot learning (ZSL) aims to classify images from unseen categories, by merely utilizing seen class images as the training data. Existing works on ZSL mainly leverage the global features or learn the global regions, from which, to construct the embeddings to the semantic space. However, few of them study the discrimination power implied in local image regions (parts), which, in some sense, correspond to semantic attributes, have stronger discrimination than attributes, and can thus assist the semantic transfer between seen/unseen classes. In this paper, to discover (semantic) regions, we propose the attentive region embedding network (AREN), which is tailored to advance the ZSL task. Specifically, AREN is end-to-end trainable and consists of two network branches, i.e., the attentive region embedding (ARE) stream, and the attentive compressed second-order embedding (ACSE) stream. ARE is capable of discovering multiple part regions under the guidance of the attention and the compatibility loss. Moreover, a novel adaptive thresholding mechanism is proposed for suppressing redundant (such as background) attention regions. To further guarantee more stable semantic transfer from the perspective of second-order collaboration, ACSE is incorporated into the AREN. In the comprehensive evaluations on four benchmarks, our models achieve state-of-the-art performances under ZSL setting, and compelling results under generalized ZSL setting. Guosen Xie, Li Liu 0004, Xiao-Bo Jin, Fan Zhu 0001, Zheng Zhang 0006, Jie Qin 0004, Yazhou Yao, Ling Shao 0001 |
CVPR | 2 |
| 2019 | Collaborative Learning of Semi-Supervised Segmentation and Classification for Medical ImagesabstractMedical image analysis has two important research areas: disease grading and fine-grained lesion segmentation. Although the former problem often relies on the latter, the two are usually studied separately. Disease severity grading can be treated as a classification problem, which only requires image-level annotations, while the lesion segmentation requires stronger pixel-level annotations. However, pixel-wise data annotation for medical images is highly time-consuming and requires domain experts. In this paper, we propose a collaborative learning method to jointly improve the performance of disease grading and lesion segmentation by semi-supervised learning with an attention mechanism. Given a small set of pixel-level annotated data, a multi-lesion mask generation model first performs the traditional semantic segmentation task. Then, based on initially predicted lesion maps for large quantities of image-level annotated data, a lesion attentive disease grading model is designed to improve the severity classification accuracy. Meanwhile, the lesion attention model can refine the lesion maps using class-specific information to fine-tune the segmentation model in a semi-supervised manner. An adversarial architecture is also integrated for training. With extensive experiments on a representative medical problem called diabetic retinopathy (DR), we validate the effectiveness of our method and achieve consistent improvements over state-of-the-art methods on three public datasets. Yi Zhou 0007, Xiaodong He 0004, Lei Huang 0015, Li Liu 0004, Fan Zhu 0001, Shanshan Cui, Ling Shao 0001 |
CVPR | 4 |
| 2019 | RANet: Ranking Attention Network for Fast Video Object SegmentationabstractDespite online learning (OL) techniques have boosted the performance of semi-supervised video object segmentation (VOS) methods, the huge time costs of OL greatly restricts their practicality. Matching based and propagation based methods run at a faster speed by avoiding OL techniques. However, they are limited by sub-optimal accuracy, due to mismatching and drifting problems. In this paper, we develop a real-time yet very accurate Ranking Attention Network (RANet) for VOS. Specifically, to integrate the insights of matching based and propagation based methods, we employ an encoder-decoder framework to learn pixel-level similarity and segmentation in an end-to-end manner. To better utilize the similarity maps, we propose a novel ranking attention module, which automatically ranks and selects these maps for fine-grained VOS performance. Experiments on DAVIS16 and DAVIS17 datasets show that our RANet achieves the best speed-accuracy trade-off, e.g., with 33 milliseconds per frame and J&F=85.5% on DAVIS16. With OL, our RANet reaches J&F=87.1% on DAVIS16, exceeding state-of-the-art VOS methods. The code can be found at https://github.com/Storife/RANet. Ziqin Wang, Jun Xu 0019, Li Liu 0004, Fan Zhu 0001, Ling Shao 0001 |
ICCV | 3 |
| 2019 | LCJoin: Set Containment Join via List CrosscuttingabstractA set containment join operates on two set-valued attributes with a subset (⊆) relationship as the join condition. It has many real-world applications, such as in publish/subscribe services and inclusion dependency discovery. Existing solutions can be broadly classified into union-oriented and intersection-oriented methods. Based on several recent studies, union-oriented methods are not competitive as they involve an expensive subset enumeration step. Intersection-oriented methods build an inverted index on one attribute and perform inverted list intersection on another attribute. Existing intersection-oriented methods intersect inverted lists one-by-one. In contrast, in this paper, we propose to intersect all the inverted lists simultaneously while skipping many irrelevant entries in the lists. To share computation, we utilize the prefix tree structure and extend our novel list intersection method to operate on the prefix tree. To further improve the efficiency, we propose to partition the data and use different methods to process each partition. We evaluated our methods using both real-world and synthetic datasets. Experimental results show that our approach outperforms existing methods by up to 10×. Dong Deng 0001, Chengcheng Yang, Shuo Shang, Fan Zhu 0001, Li Liu 0004, Ling Shao 0001 |
ICDE | 5 |
| 2019 | Toward Efficient Navigation of Massive-Scale Geo-Textual StreamsabstractWith the popularization of portable devices, numerous applications continuously produce huge streams of geo-tagged textual data, thus posing challenges to index geo-textual streaming data efficiently, which is an important task in both data management and AI applications, e.g., real-time data streams mining and targeted advertising. This, however, is not possible with the state-of-the-art indexing methods as they focus on search optimizations of static datasets, and have high index maintenance cost. In this paper, we present NQ-tree, which combines new structure designs and self-tuning methods to navigate between update and search efficiency. Our contributions include: (1) the design of multiple stores each with a different emphasis on write-friendness and read-friendness; (2) utilizing data compression techniques to reduce the I/O cost; (3) exploiting both spatial and keyword information to improve the pruning efficiency; (4) proposing an analytical cost model, and using an online self-tuning method to achieve efficient accesses to different workloads. Experiments on two real-world datasets show that NQ-tree outperforms two well designed baselines by up to 10×. Chengcheng Yang, Lisi Chen 0001, Shuo Shang, Fan Zhu 0001, Li Liu 0004, Ling Shao 0001 |
IJCAI | 5 |
| 2019 | Dynamically Visual Disambiguation of Keyword-based Image SearchabstractDue to the high cost of manual annotation, learning directly from the web has attracted broad attention. One issue that limits their performance is the problem of visual polysemy. To address this issue, we present an adaptive multi-model framework that resolves polysemy by visual disambiguation. Compared to existing methods, the primary advantage of our approach lies in that our approach can adapt to the dynamic changes in the search results. Our proposed framework consists of two major steps: we first discover and dynamically select the text queries according to the image search results, then we employ the proposed saliency-guided deep multi-instance learning network to remove outliers and learn classification models for visual disambiguation. Extensive experiments demonstrate the superiority of our proposed approach. Yazhou Yao, Zeren Sun, Fumin Shen, Li Liu 0004, Limin Wang 0002, Fan Zhu 0001, Lizhong Ding 0001, Gangshan Wu, Ling Shao 0001 |
IJCAI | 4 |
| 2019 | High-Resolution Diabetic Retinopathy Image Synthesis Manipulated by Grading and Lesions
Yi Zhou 0007, Xiaodong He 0004, Shanshan Cui, Fan Zhu 0001, Li Liu 0004, Ling Shao 0001 |
MICCAI (1) | 5 |
| 2019 | Generative Reconstructive Hashing for Incomplete Video AnalysisabstractIn the literature of video analysis, most researches, such as retrieval and recognition, hypothesize that each input video contains at least one complete semantic entity, e.g. an activity, action and event.However, this hypothesis does not hold in many realistic scenarios due to two main reasons. First, complete videos whose qualities are good enough for automatic analysis are not always accessible because of heavy motion blur, occlusions, interruptions, etc. % Second, extracting features from complete videos always fails to meet up with speed and storage requirements in large-scale use cases.To tackle these challenges, incomplete videos are more useful, but researches on them are seldom mentioned. In this paper, we propose a novel and effective hashing framework specialized in large-scale incomplete video analysis called Generative Reconstructive Hashing (GRH). To begin with, an adversarial generative network that is specially designed to map incomplete video features to the feature distributions of complete videos, so that features of incomplete videos become indistinguishable from those of complete videos. Then, the discriminative hashing module further fills the gap between full video features and estimated features from partial videos by projecting both features into a common binary feature space, which allows improvement in efficiency compared with real-value based methods. GRH is the first end-to-end framework for incomplete video analysis. Extensive experiments on various datasets demonstrate GRH's superior effectiveness and efficiency on retrieval and recognition tasks. GRH outperforms the recent state-of-the-art methods by 5.44/3.22/4.82 in terms of MAPs on HMDB51/UCF101/CCV datasets, respectively. Jingyi Zhang 0005, Zhen Wei 0001, I. C. Duta, Fumin Shen, Li Liu 0004, Fan Zhu 0001, Xing Xu 0001, Ling Shao 0001, Heng Tao Shen |
ACM Multimedia | 5 |
| 2019 | Two Generator Game: Learning to Sample via Linear Goodness-of-Fit TestabstractLearning the probability distribution of high-dimensional data is a challenging problem. To solve this problem, we formulate a deep energy adversarial network (DEAN), which casts the energy model learned from real data into an optimization of a goodness-of-fit (GOF) test statistic. DEAN can be interpreted as a GOF game between two generative networks, where one explicit generative network learns an energy-based distribution that fits the real data, and the other implicit generative network is trained by minimizing a GOF test statistic between the energy-based distribution and the generated data, such that the underlying distribution of the generated data is close to the energy-based distribution. We design a two-level alternative optimization procedure to train the explicit and implicit generative networks, such that the hyper-parameters can also be automatically learned. Experimental results show that DEAN achieves high quality generations compared to the state-of-the-art approaches. Lizhong Ding 0001, Mengyang Yu, Li Liu 0004, Fan Zhu 0001, Yong Liu 0018, Yu Li 0006, Ling Shao 0001 |
NeurIPS | 3 |
| 2019 | Unsupervised Binary Representation Learning with Deep Variational Networks
Yuming Shen, Li Liu 0004, Ling Shao 0001 |
Int. J. Comput. Vis. | 2 |
| 2019 | Clustering-driven unsupervised deep hashing for image retrieval
Haofeng Zhang 0001, Yazhou Yao, Wankou Yang, Li Liu 0004 |
Neurocomputing | 6 |
| 2019 | Adversarial unseen visual feature synthesis for Zero-shot Learning
Haofeng Zhang 0001, Yang Long 0001, Li Liu 0004, Ling Shao 0001 |
Neurocomputing | 3 |
| 2019 | Multiview discriminative marginal metric learning for makeup face verificationabstractMakeup face verification in the wild is an important research problem for its popularization in real-world. However, little effort has been made to tackle it in computer vision. In this research, we first build a new database, i.e., Facial Beauty Database (FBD), which contains paired facial images of 8933 subjects without and with makeup in different real-world scenarios. To the best of our knowledge, FBD is the largest makeup face database to date compared with existing databases for facial makeup research. Moreover, we propose a new discriminative marginal metric learning (DMML) algorithm to deal with this problem in the wild. Inspired by the fact that interclass marginal faces are usually more discriminative than interclass nonmarginal faces in learning the discriminative metric space, we use the interclass marginal faces to depict the discriminative information. Simultaneously, we wish that those interclass marginal faces without makeup relations are separated from each other as far as possible, so that more discriminative information between facial images without and with makeup can be exploited for verification. Furthermore, since multiple features could provide comprehensive information in describing the facial representations from diverse points of view and extract more informative cues from facial images, we also introduce a multiview discriminative marginal metric learning (MDMML) algorithm by effectively learning a robust metric space such that multiple features from different points of view can be integrated to effectively enhance the performance of makeup face verification. Experimental results on two real-world makeup face databases are utilized to show the effectiveness of our method and the possibility of verifying the makeup relations from facial images in real-world. Lining Zhang, Hubert P. H. Shum, Li Liu 0004, Guodong Guo, Ling Shao 0001 |
Neurocomputing | 3 |
| 2019 | Binary Multi-View ClusteringabstractClustering is a long-standing important research problem, however, remains challenging when handling large-scale image data from diverse sources. In this paper, we present a novel Binary Multi-View Clustering (BMVC) framework, which can dexterously manipulate multi-view image data and easily scale to large data. To achieve this goal, we formulate BMVC by two key components: compact collaborative discrete representation learning and binary clustering structure learning, in a joint learning framework. Specifically, BMVC collaboratively encodes the multi-view image descriptors into a compact common binary code space by considering their complementary information; the collaborative binary representations are meanwhile clustered by a binary matrix factorization model, such that the cluster structures are optimized in the Hamming space by pure, extremely fast bit-operations. For efficiency, the code balance constraints are imposed on both binary data representations and cluster centroids. Finally, the resulting optimization problem is solved by an alternating optimization scheme with guaranteed fast convergence. Extensive experiments on four large-scale multi-view image datasets demonstrate that the proposed method enjoys the significant reduction in both computation and memory footprint, while observing superior (in most cases) or very competitive performance, in comparison with state-of-the-art clustering methods. Zheng Zhang 0006, Li Liu 0004, Fumin Shen, Heng Tao Shen, Ling Shao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2019 | Deep Self-Taught Hashing for Image RetrievalabstractHashing algorithm has been widely used to speed up image retrieval due to its compact binary code and fast distance calculation. The combination with deep learning boosts the performance of hashing by learning accurate representations and complicated hashing functions. So far, the most striking success in deep hashing have mostly involved discriminative models, which require labels. To apply deep hashing on datasets without labels, we propose a deep self-taught hashing algorithm (DSTH), which generates a set of pseudo labels by analyzing the data itself, and then learns the hash functions for novel data using discriminative deep models. Furthermore, we generalize DSTH to support both supervised and unsupervised cases by adaptively incorporating label information. We use two different deep learning framework to train the hash functions to deal with out-of-sample problem and reduce the time complexity without loss of accuracy. We have conducted extensive experiments to investigate different settings of DSTH, and compared it with state-of-the-art counterparts in six publicly available datasets. The experimental results show that DSTH outperforms the others in all datasets. Yu Liu 0040, Jingkuan Song, Ke Zhou 0001, Lingyu Yan, Li Liu 0004, Fuhao Zou, Ling Shao 0001 |
IEEE Trans. Cybern. | 5 |
| 2019 | Scalable Zero-Shot Learning via Binary Visual-Semantic EmbeddingsabstractZero-shot learning aims to classify visual instances from unseen classes in the absence of training examples. This is typically achieved by directly mapping visual features to a semantic embedding space of classes (e.g., attributes or word vectors), where the similarity between the two modalities can be readily measured. However, the semantic space may not be reliable for recognition due to the noisy class embeddings or visual bias problem. In this work, we propose a novel Binary embedding based Zero-Shot Learning (BZSL) method, which recognizes visual instances from unseen classes through an intermediate discriminative Hamming space. Specifically, BZSL jointly learns two binary coding functions to encode both visual instances and class embeddings into the Hamming space, which well alleviates the visual-semantic bias problem. As a desiring property, classifying an unseen instance thereby can be efficiently done by retrieving its nearest-class codes with minimal Hamming distance. During training, by introducing two auxiliary variables for the coding functions, we formulate an equivalent correlation maximization problem, which admits an analytical solution. The resulting algorithm thus enjoys both highly efficient training and scalable novel class inferring. Extensive experiments on four benchmark datasets, including the full ImageNet Fall 2011 dataset with over 20K unseen classes, demonstrate the superiority of our method on the zero-shot learning task. Particularly, we show that increasing the binary embedding dimension can inevitably improve the recognition accuracy. Fumin Shen, Xiang Zhou 0008, Jun Yu 0002, Yang Yang 0002, Li Liu 0004, Heng Tao Shen |
IEEE Trans. Image Process. | 5 |
| 2019 | Unsupervised Deep Video Hashing via Balanced Code for Large-Scale Video RetrievalabstractThis paper proposes a deep hashing framework, namely Unsupervised Deep Video Hashing (UDVH), for largescale video similarity search with the aim to learn compact yet effective binary codes. Our UDVH produces the hash codes in a self-taught manner by jointly integrating discriminative video representation with optimal code learning, where an efficient alternating approach is adopted to optimize the objective function. The key differences from most existing video hashing methods lie in 1) UDVH is an unsupervised hashing method that generates hash codes by cooperatively utilizing feature clustering and a specifically-designed binarization with the original neighborhood structure preserved in the binary space; 2) a specific rotation is developed and applied onto video features such that the variance of each dimension can be balanced, thus facilitating the subsequent quantization step. Extensive experiments performed on three popular video datasets show that UDVH is overwhelmingly better than the state-of-the-arts in terms of various evaluation metrics, which makes it practical in real-world applications. Gengshen Wu, Jungong Han, Li Liu 0004, Guiguang Ding, Qiang Ni, Ling Shao 0001 |
IEEE Trans. Image Process. | 4 |
| 2019 | Extracting Privileged Information for Enhancing Classifier LearningabstractThe accuracy of data-driven learning approaches is often unsatisfactory when the training data is inadequate either in quantity or quality. Manually labeled privileged information (PI), e.g., attributes, tags or properties, is usually incorporated to improve classifier learning. However, the process of manually labeling is time-consuming and labor-intensive. Moreover, due to the limitations of personal knowledge, manually labeled PI may not be rich enough. To address these issues, we propose to enhance classifier learning by exploring PI from untagged corpora, which can effectively eliminate the dependency on manually labeled data and obtain much richer PI. In detail, we treat each selected PI as a subcategory and learn one classifier for each subcategory independently. The classifiers for all subcategories are integrated together to form a more powerful category classifier. Particularly, we propose a novel instancelevel multi-instance learning (MIL) model to simultaneously select a subset of training images from each subcategory and learn the optimal SVM classifiers based on the selected images. Extensive experiments on four benchmark datasets demonstrate the superiority of our proposed approach. Yazhou Yao, Fumin Shen, Jian Zhang 0002, Li Liu 0004, Zhenmin Tang, Ling Shao 0001 |
IEEE Trans. Image Process. | 4 |
| 2019 | Scalable Supervised Asymmetric Hashing With Semantic and Latent Factor EmbeddingabstractCompact hash code learning has been widely applied to fast similarity search owing to its significantly reduced storage and highly efficient query speed. However, it is still a challenging task to learn discriminative binary codes for perfectly preserving the full pairwise similarities embedded in the high-dimensional real-valued features, such that the promising performance can be guaranteed. To overcome this difficulty, in this paper, we propose a novel scalable supervised asymmetric hashing (SSAH) method, which can skillfully approximate the full-pairwise similarity matrix based on maximum asymmetric inner product of two different non-binary embeddings. In particular, to comprehensively explore the semantic information of data, the supervised label information and the refined latent feature embedding are simultaneously considered to construct the high-quality hashing function and boost the discriminant of the learned binary codes. Specifically, SSAH learns two distinctive hashing functions in conjunction of minimizing the regression loss on the semantic label alignment and the encoding loss on the refined latent features. More importantly, instead of using only part of similarity correlations of data, the full-pairwise similarity matrix is directly utilized to avoid information loss and performance degeneration, and its cumbersome computation complexity on n ×n matrix can be dexterously manipulated during the optimization phase. Furthermore, an efficient alternating optimization scheme with guaranteed convergence is designed to address the resulting discrete optimization problem. The encouraging experimental results on diverse benchmark datasets demonstrate the superiority of the proposed SSAH method in comparison with many recently proposed hashing algorithms. Zheng Zhang 0006, Zhihui Lai 0001, Zi Huang, Wai Keung Wong, Guosen Xie, Li Liu 0004, Ling Shao 0001 |
IEEE Trans. Image Process. | 6 |
| 2019 | Extracting Multiple Visual Senses for Web LearningabstractLabeled image datasets have played a critical role in high-level image understanding. However, the process of manual labeling is both time consuming and labor intensive. To reduce the dependence on manually labeled data, there have been increasing research efforts on learning visual classifiers by directly exploiting web images. One issue that limits their performance is the problem of polysemy. Existing unsupervised approaches attempt to reduce the influence of visual polysemy by filtering out irrelevant images, but do not directly address polysemy. To this end, in this paper, we present a multimodal framework that solves the problem of polysemy by allowing sense-specific diversity in search results. Specifically, we first discover a list of possible semantic senses from untagged corpora to retrieve sense-specific images. Then, we merge visual similar semantic senses and prune noise by using the retrieved images. Finally, we train one visual classifier for each selected semantic sense and use the learned sense-specific classifiers to distinguish multiple visual senses. Extensive experiments on classifying images into sense-specific categories and reranking search results demonstrate the superiority of our proposed approach. Yazhou Yao, Fumin Shen, Jian Zhang 0002, Li Liu 0004, Zhenmin Tang, Ling Shao 0001 |
IEEE Trans. Multim. | 4 |
| 2018 | Towards Affordable Semantic Searching: Zero-Shot Retrieval via Dominant AttributesabstractInstance-level retrieval has become an essential paradigm to index and retrieves images from large-scale databases. Conventional instance search requires at least an example of the query image to retrieve images that contain the same object instance. Existing semantic retrieval can only search semantically-related images, such as those sharing the same category or a set of tags, not the exact instances. Meanwhile, the unrealistic assumption is that all categories or tags are known beforehand. Training models for these semantic concepts highly rely on instance-level attributes or human captions which are expensive to acquire. Given the above challenges, this paper studies the Zero-shot Retrieval problem that aims for instance-level image search using only a few dominant attributes. The contributions are: 1) we utilise automatic word embedding to infer class-level attributes to circumvent expensive human labelling; 2) the inferred class-attributes can be extended into discriminative instance attributes through our proposed Latent Instance Attributes Discovery (LIAD) algorithm; 3) our method is not restricted to complete attribute signatures, query of dominant attributes can also be dealt with. On two benchmarks, CUB and SUN, extensive experiments demonstrate that our method can achieve promising performance for the problem. Moreover, our approach can also benefit conventional ZSL tasks. Yang Long 0001, Li Liu 0004, Yuming Shen, Ling Shao 0001 |
AAAI | 2 |
| 2018 | Structure-Aware 3D Shape Synthesis from Single-View Images
Xuyang Hu, Fan Zhu 0001, Li Liu 0004, Jin Xie 0001, Jun Tang 0007, Nian Wang 0002, Fumin Shen, Ling Shao 0001 |
BMVC | 3 |
| 2018 | Pixel-level Semantics Guided Image Colorization
Jiaojiao Zhao, Li Liu 0004, Cees Snoek, Jungong Han, Ling Shao 0001 |
BMVC | 2 |
| 2018 | Zero-Shot Sketch-Image HashingabstractRecent studies show that large-scale sketch-based image retrieval (SBIR) can be efficiently tackled by cross-modal binary representation learning methods, where Hamming distance matching significantly speeds up the process of similarity search. Providing training and test data subjected to a fixed set of pre-defined categories, the cutting-edge SBIR and cross-modal hashing works obtain acceptable retrieval performance. However, most of the existing methods fail when the categories of query sketches have never been seen during training. In this paper, the above problem is briefed as a novel but realistic zero-shot SBIR hashing task. We elaborate the challenges of this special task and accordingly propose a zero-shot sketch-image hashing (ZSIH) model. An end-to-end three-network architecture is built, two of which are treated as the binary encoders. The third network mitigates the sketch-image heterogeneity and enhances the semantic relations among data by utilizing the Kronecker fusion layer and graph convolution, respectively. As an important part of ZSIH, we formulate a generative hashing scheme in reconstructing semantic knowledge representations for zero-shot retrieval. To the best of our knowledge, ZSIH is the first zero-shot hashing work suitable for SBIR and cross-modal search. Comprehensive experiments are conducted on two extended datasets, i.e., Sketchy and TU-Berlin with a novel zero-shot train-test split. The proposed model remarkably outperforms related works. Yuming Shen, Li Liu 0004, Fumin Shen, Ling Shao 0001 |
CVPR | 2 |
| 2018 | Deep Multi-task Learning to Recognise Subtle Facial Expressions of Mental States
Guosheng Hu, Li Liu 0004, Yang Hua 0001, Zhihong Zhang 0001, Fumin Shen, Ling Shao 0001, Timothy M. Hospedales, Neil Robertson 0002, Yongxin Yang |
ECCV (12) | 2 |
| 2018 | TBN: Convolutional Neural Network with Ternary Inputs and Binary Weights
Diwen Wan, Fumin Shen, Li Liu 0004, Fan Zhu 0001, Jie Qin 0004, Ling Shao 0001, Heng Tao Shen |
ECCV (2) | 3 |
| 2018 | Highly-Economized Multi-view Binary Compression for Scalable Image Clustering
Zheng Zhang 0006, Li Liu 0004, Jie Qin 0004, Fan Zhu 0001, Fumin Shen, Yong Xu 0001, Ling Shao 0001, Heng Tao Shen |
ECCV (12) | 2 |
| 2018 | Generative Domain-Migration Hashing for Sketch-to-Image Retrieval
Jingyi Zhang 0005, Fumin Shen, Li Liu 0004, Fan Zhu 0001, Mengyang Yu, Ling Shao 0001, Heng Tao Shen, Luc Van Gool |
ECCV (2) | 3 |
| 2018 | Unsupervised Deep Hashing via Binary Latent Factor Models for Large-scale Cross-modal RetrievalabstractDespite its great success, matrix factorization based cross-modality hashing suffers from two problems: 1) there is no engagement between feature learning and binarization; and 2) most existing methods impose the relaxation strategy by discarding the discrete constraints when learning the hash function, which usually yields suboptimal solutions. In this paper, we propose a novel multimodal hashing framework, referred as Unsupervised Deep Cross-Modal Hashing (UDCMH), for multimodal data search in a self-taught manner via integrating deep learning and matrix factorization with binary latent factor models. On one hand, our unsupervised deep learning framework enables the feature learning to be jointly optimized with the binarization. On the other hand, the hashing system based on the binary latent factor models can generate unified binary codes by solving a discrete-constrained objective function directly with no need for a relaxation step. Moreover, novel Laplacian constraints are incorporated into the objective function, which allow to preserve not only the nearest neighbors that are commonly considered in the literature but also the farthest neighbors of data, even if the semantic labels are not available. Extensive experiments on multiple datasets highlight the superiority of the proposed framework over several state-of-the-art baselines. Gengshen Wu, Zijia Lin, Jungong Han, Li Liu 0004, Guiguang Ding, Baochang Zhang 0001, Jialie Shen 0001 |
IJCAI | 4 |
| 2018 | Learning to Synthesize 3D Indoor Scenes from Monocular ImagesabstractDepth images have always been playing critical roles for indoor scene understanding problems, and are particularly important for tasks in which 3D inferences are involved. However, since depth images are not universally available, abandoning them from the testing stage can significantly improve the generality of a method. In this work, we consider the scenarios where depth images are not available in the testing data, and propose to learn a convolutional long short-term memory (Conv LSTM) network and a regression convolutional neural network (regression ConvNet) using only monocular RGB images. The proposed networks benefit from 2D segmentations, object-level spatial context, object-scene dependencies and objects' geometric information, where optimization is governed by the semantic label loss, which measures the label consistencies of both objects and scenes, and the 3D geometrical loss, which measures the correctness of objects' 6Dof estimation. Conv LSTM and regression ConvNet are applied to scene/object classification, object detection and 6Dof estimation tasks respectively, where we utilize the joint inference from both networks and further provide the perspective of synthesizing fully rigged 3D scenes according to objects' arrangements in monocular images. Both quantitative and qualitative experimental results are provided on the NYU-v2 dataset, and we demonstrate that the proposed Conv LSTM can achieve state-of-the-art performance without requiring the depth information. Fan Zhu 0001, Li Liu 0004, Jin Xie 0001, Fumin Shen, Ling Shao 0001, Yi Fang 0006 |
ACM Multimedia | 2 |
| 2018 | Visual Spatial Attention Network for Relationship DetectionabstractVisual relationship detection, which aims to predict a triplet with the detected objects, has attracted increasing attention in the scene understanding study. During tackling this problem, dealing with varying scales of the subjects and objects is of great importance, which has been less studied. To overcome this challenge, we propose a novel Vision Spatial Attention Network (VSA-Net), which employs a two-dimensional normal distribution attention scheme to effectively model small objects. In addition, we design a Subject-Object-layer (SO-layer) to distinguish between the subject and object to attain more precise results. To the best of our knowledge, VSA-Net is the first end-to-end attention mechanism based visual relationship detection model. Extensive experiments on the benchmark datasets (VRD and VG) show that, by using pure vision information, our VSA-Net achieves state-of-the-art performance for predicate detection, phrase detection, and relationship detection. Chaojun Han, Fumin Shen, Li Liu 0004, Yang Yang 0002, Heng Tao Shen |
ACM Multimedia | 3 |
| 2018 | GraphCAR: Content-aware Multimedia Recommendation with Graph AutoencoderabstractPrecisely recommending relevant multimedia items from massive candidates to a large number of users is an indispensable yet difficult task on many platforms. A promising way is to project users and items into a latent space and recommend items via the inner product of latent factor vectors. However, previous studies paid little attention to the multimedia content itself and couldn't make the best use of preference data like implicit feedback. To fill this gap, we propose a Content-aware Multimedia Recommendation Model with Graph Autoencoder (GraphCAR), combining informative multimedia content with user-item interaction. Specifically, user-item interaction, user attributes and multimedia contents (e.g., images, videos, audios, etc.) are taken as input of the autoencoder to generate the item preference scores for each user. Through extensive experiments on two real-world multimedia Web services: Amazon and Vine, we show that GraphCAR significantly outperforms state-of-the-art techniques of both collaborative filtering and content-based methods. Qidi Xu, Fumin Shen, Li Liu 0004, Heng Tao Shen |
SIGIR | 3 |
| 2018 | Zero-Shot Learning Using Synthesised Unseen Visual Data with Diffusion RegularisationabstractSufficient training examples are the fundamental requirement for most of the learning tasks. However, collecting well-labelled training examples is costly. Inspired by Zero-shot Learning (ZSL) that can make use of visual attributes or natural language semantics as an intermediate level clue to associate low-level features with high-level classes, in a novel extension of this idea, we aim to synthesise training data for novel classes using only semantic attributes. Despite the simplicity of this idea, there are several challenges. First, how to prevent the synthesised data from over-fitting to training classes? Second, how to guarantee the synthesised data is discriminative for ZSL tasks? Third, we observe that only a few dimensions of the learnt features gain high variances whereas most of the remaining dimensions are not informative. Thus, the question is how to make the concentrated information diffuse to most of the dimensions of synthesised data. To address the above issues, we propose a novel embedding algorithm named Unseen Visual Data Synthesis (UVDS) that projects semantic features to the high-dimensional visual feature space. Two main techniques are introduced in our proposed algorithm. (1) We introduce a latent embedding space which aims to reconcile the structural difference between the visual and semantic spaces, meanwhile preserve the local structure. (2) We propose a novel Diffusion Regularisation (DR) that explicitly forces the variances to diffuse over most dimensions of the synthesised data. By an orthogonal rotation (more precisely, an orthogonal transformation), DR can remove the redundant correlated attributes and further alleviate the over-fitting problem. On four benchmark datasets, we demonstrate the benefit of using synthesised unseen data for zero-shot learning. Extensive experimental results suggest that our proposed approach significantly outperforms the state-of-the-art methods. Yang Long 0001, Li Liu 0004, Fumin Shen, Ling Shao 0001, Xuelong Li 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2018 | Unsupervised Deep Hashing with Similarity-Adaptive and Discrete OptimizationabstractRecent vision and learning studies show that learning compact hash codes can facilitate massive data processing with significantly reduced storage and computation. Particularly, learning deep hash functions has greatly improved the retrieval performance, typically under the semantic supervision. In contrast, current unsupervised deep hashing algorithms can hardly achieve satisfactory performance due to either the relaxed optimization or absence of similarity-sensitive objective. In this work, we propose a simple yet effective unsupervised hashing framework, named Similarity-Adaptive Deep Hashing (SADH), which alternatingly proceeds over three training modules: deep hash model training, similarity graph updating and binary code optimization. The key difference from the widely-used two-step hashing method is that the output representations of the learned deep model help update the similarity graph matrix, which is then used to improve the subsequent code optimization. In addition, for producing high-quality binary codes, we devise an effective discrete optimization algorithm which can directly handle the binary constraints with a general hashing loss. Extensive experiments validate the efficacy of SADH, which consistently outperforms the state-of-the-arts by large gaps. Fumin Shen, Yan Xu 0009, Li Liu 0004, Yang Yang 0002, Zi Huang, Heng Tao Shen |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2018 | Quantization-based hashing: a general framework for scalable image and video retrieval
Jingkuan Song, Lianli Gao, Li Liu 0004, Xiaofeng Zhu 0001, Nicu Sebe |
Pattern Recognit. | 3 |
| 2018 | Zero-shot learning via discriminative representation extraction
Teng Long 0002, Xing Xu 0001, Fumin Shen, Li Liu 0004, Ning Xie 0003, Yang Yang 0002 |
Pattern Recognit. Lett. | 4 |
| 2018 | Dense Invariant Feature-Based Support Vector Ranking for Cross-Camera Person ReidentificationabstractRecently, support vector ranking (SVR) has been adopted to address the challenging person reidentification problem. However, the ranking model based on ordinary global features cannot well represent the significant variation of pose and viewpoint across camera views. To address this issue, a novel ranking method that fuses the dense invariant features (DIFs) is proposed in this paper to model the variation of images across camera views. An optimal space for ranking is learned by simultaneously maximizing the margin and minimizing the error on the fused features. The proposed method significantly outperforms the original SVR algorithm due to the invariance of the DIFs, the fusion of the bidirectional features, and the adaptive adjustment of parameters. Experimental results demonstrate that the proposed method is competitive with state-of-the-art methods on two challenging data sets, showing its potential for real-world person reidentification. Shoubiao Tan, Feng Zheng 0001, Li Liu 0004, Jungong Han, Ling Shao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2018 | Deep Action Parsing in Videos With Large-Scale Synthesized DataabstractAction parsing in videos with complex scenes is an interesting but challenging task in computer vision. In this paper, we propose a generic 3D convolutional neural network in a multi-task learning manner for effective Deep Action Parsing (DAP3D-Net) in videos. Particularly, in the training phase, action localization, classification, and attributes learning can be jointly optimized on our appearance-motion data via DAP3D-Net. For an upcoming test video, we can describe each individual action in the video simultaneously as: Where the action occurs, What the action is, and How the action is performed. To well demonstrate the effectiveness of the proposed DAP3D-Net, we also contribute a new Numerous-category Aligned Synthetic Action data set, i.e., NASA, which consists of 200 000 action clips of over 300 categories and with 33 pre-defined action attributes in two hierarchical levels (i.e., low-level attributes of basic body part movements and high-level attributes related to action motion). We learn DAP3D-Net using the NASA data set and then evaluate it on our collected Human Action Understanding data set and the public THUMOS data set. Experimental results show that our approach can accurately localize, categorize, and describe multiple actions in realistic videos. Li Liu 0004, Yi Zhou 0007, Ling Shao 0001 |
IEEE Trans. Image Process. | 1 |
| 2018 | Unsupervised Deep Hashing With Pseudo Labels for Scalable Image RetrievalabstractIn order to achieve efficient similarity searching, hash functions are designed to encode images into low-dimensional binary codes with the constraint that similar features will have a short distance in the projected Hamming space. Recently, deep learning-based methods have become more popular, and outperform traditional non-deep methods. However, without label information, most state-of-the-art unsupervised deep hashing (DH) algorithms suffer from severe performance degradation for unsupervised scenarios. One of the main reasons is that the ad-hoc encoding process cannot properly capture the visual feature distribution. In this paper, we propose a novel unsupervised framework that has two main contributions: 1) we convert the unsupervised DH model into supervised by discovering pseudo labels; 2) the framework unifies likelihood maximization, mutual information maximization, and quantization error minimization so that the pseudo labels can maximumly preserve the distribution of visual features. Extensive experiments on three popular data sets demonstrate the advantages of the proposed method, which leads to significant performance improvement over the state-of-the-art unsupervised hashing algorithms. Haofeng Zhang 0001, Li Liu 0004, Yang Long 0001, Ling Shao 0001 |
IEEE Trans. Image Process. | 2 |
| 2018 | Vehicle Re-Identification by Deep Hidden Multi-View InferenceabstractVehicle re-identification (re-ID) is an area that has received far less attention in the computer vision community than the prevalent person re-ID. Possible reasons for this slow progress are the lack of appropriate research data and the special 3D structure of a vehicle. Previous works have generally focused on some specific views (e.g., front); but, these methods are less effective in realistic scenarios, where vehicles usually appear in arbitrary views to cameras. In this paper, we focus on the uncertainty of vehicle viewpoint in re-ID, proposing two end-to-end deep architectures: the Spatially Concatenated ConvNet and convolutional neural network (CNN)-LSTM bi-directional loop. Our models exploit the great advantages of the CNN and long short-term memory (LSTM) to learn transformations across different viewpoints of vehicles. Thus, a multi-view vehicle representation containing all viewpoints' information can be inferred from the only one input view, and then used for learning to measure distance. To verify our models, we also introduce a Toy Car RE-ID data set with images from multiple viewpoints of 200 vehicles. We evaluate our proposed methods on the Toy Car RE-ID data set and the public Multi-View Car, VehicleID, and VeRi data sets. Experimental results illustrate that our models achieve consistent improvements over the state-of-the-art vehicle re-ID approaches. Yi Zhou 0007, Li Liu 0004, Ling Shao 0001 |
IEEE Trans. Image Process. | 2 |
| 2018 | Multiview Discrete Hashing for Scalable Multimedia SearchabstractHashing techniques have recently gained increasing research interest in multimedia studies. Most existing hashing methods only employ single features for hash code learning. Multiview data with each view corresponding to a type of feature generally provides more comprehensive information. How to efficiently integrate multiple views for learning compact hash codes still remains challenging. In this article, we propose a novel unsupervised hashing method, dubbed multiview discrete hashing (MvDH), by effectively exploring multiview data. Specifically, MvDH performs matrix factorization to generate the hash codes as the latent representations shared by multiple views, during which spectral clustering is performed simultaneously. The joint learning of hash codes and cluster labels enables that MvDH can generate more discriminative hash codes, which are optimal for classification. An efficient alternating algorithm is developed to solve the proposed optimization problem with guaranteed convergence and low computational complexity. The binary codes are optimized via the discrete cyclic coordinate descent (DCC) method to reduce the quantization errors. Extensive experimental results on three large-scale benchmark datasets demonstrate the superiority of the proposed method over several state-of-the-art methods in terms of both accuracy and scalability. Xiaobo Shen 0001, Fumin Shen, Li Liu 0004, Yun-Hao Yuan 0001, Weiwei Liu 0003, Quan-Sen Sun |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2018 | Fast Automatic Vehicle Annotation for Urban Traffic SurveillanceabstractAutomatic vehicle detection and annotation for streaming video data with complex scenes is an interesting but challenging task for intelligent transportation systems. In this paper, we present a fast algorithm: detection and annotation for vehicles (DAVE), which effectively combines vehicle detection and attributes annotation into a unified framework. DAVE consists of two convolutional neural networks: a shallow fully convolutional fast vehicle proposal network (FVPN) for extracting all vehicles' positions, and a deep attributes learning network (ALN), which aims to verify each detection candidate and infer each vehicle's pose, color, and type information simultaneously. These two nets are jointly optimized so that abundant latent knowledge learned from the deep empirical ALN can be exploited to guide training the much simpler FVPN. Once the system is trained, DAVE can achieve efficient vehicle detection and attributes annotation for real-world traffic surveillance data, while the FVPN can be independently adopted as a real-time high-performance vehicle detector as well. We evaluate the DAVE on a new self-collected urban traffic surveillance data set and the public PASCAL VOC2007 car and LISA 2010 data sets, with consistent improvements over existing algorithms. Yi Zhou 0007, Li Liu 0004, Ling Shao 0001, Matt Mellor |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2018 | Marginal Representation Learning With Graph Structure Self-AdaptationabstractLearning discriminative feature representations has shown remarkable importance due to its promising performance for machine learning problems. This paper presents a discriminative data representation learning framework by employing a simple yet powerful marginal regression function with probabilistic graphical structure adaptation. A marginally structured representation learning (MSRL) method is proposed by seamlessly incorporating distinguishable regression targets analysis, graph structure adaptation, and robust linear structural learning into a joint framework. Specifically, MSRL learns marginal regression targets from data rather than exploiting the conventional zero-one matrix that greatly hinders the freedom of regression fitness and degrades the performance of regression results. Meanwhile, an optimized graph regularization term with self-improving adaptation is constructed based on probabilistic connection knowledge to improve the compactness of the learned representation. Additionally, the regression targets are further predicted by utilizing the explanatory factors from the latent subspace of data, which can uncover the underlying feature correlations to enhance the reliability. The resulting optimization problem can be elegantly solved by an efficient iterative algorithm. Finally, the proposed method is evaluated by eight diverse but related tasks, including object, face, texture, and scene, categorization data sets. The encouraging experimental results and the explicit theoretical analysis demonstrate the efficacy of the proposed representation learning method in comparison with state-of-the-art algorithms. Zheng Zhang 0006, Ling Shao 0001, Yong Xu 0001, Li Liu 0004, Jian Yang 0003 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2017 | Unsupervised Deep Generative Hashing
Yuming Shen, Li Liu 0004, Ling Shao 0001 |
BMVC | 2 |
| 2017 | Fast Person Re-identification via Cross-Camera Semantic Binary TransformationabstractNumerous methods have been proposed for person re-identification, most of which however neglect the matching efficiency. Recently, several hashing based approaches have been developed to make re-identification more scalable for large-scale gallery sets. Despite their efficiency, these works ignore cross-camera variations, which severely deteriorate the final matching accuracy. To address the above issues, we propose a novel hashing based method for fast person re-identification, namely Cross-camera Semantic Binary Transformation (CSBT). CSBT aims to transform original high-dimensional feature vectors into compact identity-preserving binary codes. To this end, CSBT first employs a subspace projection to mitigate cross-camera variations, by maximizing intra-person similarities and inter-person discrepancies. Subsequently, a binary coding scheme is proposed via seamlessly incorporating both the semantic pairwise relationships and local affinity information. Finally, a joint learning framework is proposed for simultaneous subspace projection learning and binary coding based on discrete alternating optimization. Experimental results on four benchmarks clearly demonstrate the superiority of CSBT over the state-of-the-art methods. Jiaxin Chen 0002, Yunhong Wang 0001, Jie Qin 0004, Li Liu 0004, Ling Shao 0001 |
CVPR | 4 |
| 2017 | Deep Sketch Hashing: Fast Free-Hand Sketch-Based Image RetrievalabstractFree-hand sketch-based image retrieval (SBIR) is a specific cross-view retrieval task, in which queries are abstract and ambiguous sketches while the retrieval database is formed with natural images. Work in this area mainly focuses on extracting representative and shared features for sketches and natural images. However, these can neither cope well with the geometric distortion between sketches and images nor be feasible for large-scale SBIR due to the heavy continuous-valued distance computation. In this paper, we speed up SBIR by introducing a novel binary coding method, named Deep Sketch Hashing (DSH), where a semi-heterogeneous deep architecture is proposed and incorporated into an end-to-end binary coding framework. Specifically, three convolutional neural networks are utilized to encode free-hand sketches, natural images and, especially, the auxiliary sketch-tokens which are adopted as bridges to mitigate the sketch-image geometric distortion. The learned DSH codes can effectively capture the cross-view similarities as well as the intrinsic semantic correlations between different categories. To the best of our knowledge, DSH is the first hashing work specifically designed for category-level SBIR with an end-to-end deep architecture. The proposed DSH is comprehensively evaluated on two large-scale datasets of TU-Berlin Extension and Sketchy, and the experiments consistently show DSHs superior SBIR accuracies over several state-of-the-art methods, while achieving significantly reduced retrieval time and memory footprint. Li Liu 0004, Fumin Shen, Yuming Shen, Xianglong Liu 0001, Ling Shao 0001 |
CVPR | 1 |
| 2017 | Discretely Coding Semantic Rank Orders for Supervised Image HashingabstractLearning to hash has been recognized to accomplish highly efficient storage and retrieval for large-scale visual data. Particularly, ranking-based hashing techniques have recently attracted broad research attention because ranking accuracy among the retrieved data is well explored and their objective is more applicable to realistic search tasks. However, directly optimizing discrete hash codes without continuous-relaxations on a nonlinear ranking objective is infeasible by either traditional optimization methods or even recent discrete hashing algorithms. To address this challenging issue, in this paper, we introduce a novel supervised hashing method, dubbed Discrete Semantic Ranking Hashing (DSeRH), which aims to directly embed semantic rank orders into binary codes. In DSeRH, a generalized Adaptive Discrete Minimization (ADM) approach is proposed to discretely optimize binary codes with the quadratic nonlinear ranking objective in an iterative manner and is guaranteed to converge quickly. Additionally, instead of using 0/1 independent labels to form rank orders as in previous works, we generate the listwise rank orders from the high-level semantic word embeddings which can quantitatively capture the intrinsic correlation between different categories. We evaluate our DSeRH, coupled with both linear and deep convolutional neural network (CNN) hash functions, on three image datasets, i.e., CIFAR-10, SUN397 and ImageNet100, and the results manifest that DSeRH can outperform the state-of-the-art ranking-based hashing methods. Li Liu 0004, Ling Shao 0001, Fumin Shen, Mengyang Yu |
CVPR | 1 |
| 2017 | From Zero-Shot Learning to Conventional Supervised Classification: Unseen Visual Data SynthesisabstractRobust object recognition systems usually rely on powerful feature extraction mechanisms from a large number of real images. However, in many realistic applications, collecting sufficient images for ever-growing new classes is unattainable. In this paper, we propose a new Zero-shot learning (ZSL) framework that can synthesise visual features for unseen classes without acquiring real images. Using the proposed Unseen Visual Data Synthesis (UVDS) algorithm, semantic attributes are effectively utilised as an intermediate clue to synthesise unseen visual features at the training stage. Hereafter, ZSL recognition is converted into the conventional supervised problem, i.e. the synthesised visual features can be straightforwardly fed to typical classifiers such as SVM. On four benchmark datasets, we demonstrate the benefit of using synthesised unseen data. Extensive experimental results manifest that our proposed approach significantly improve the state-of-the-art results. Yang Long 0001, Li Liu 0004, Ling Shao 0001, Fumin Shen, Guiguang Ding, Jungong Han |
CVPR | 2 |
| 2017 | Binary Coding for Partial Action Analysis with Limited Observation RatiosabstractTraditional action recognition methods aim to recognize actions with complete observations/executions. However, it is often difficult to capture fully executed actions due to occlusions, interruptions, etc. Meanwhile, action prediction/recognition in advance based on partial observations is essential for preventing the situation from deteriorating. Besides, fast spotting human activities using partially observed data is a critical ingredient for retrieval systems. Inspired by the recent success of data binarization in efficient retrieval/recognition, we propose a novel approach, named Partial Reconstructive Binary Coding (PRBC), for action analysis based on limited frame glimpses during any period of the complete execution. Specifically, we learn discriminative compact binary codes for partial actions via a joint learning framework, which collaboratively tackles feature reconstruction as well as binary coding. We obtain the solution to PRBC based on a discrete alternating iteration algorithm. Extensive experiments on four realistic action datasets in terms of three tasks (i.e., partial action retrieval, recognition and prediction) clearly show the superiority of PRBC over the state-of-the-art methods, along with significantly reduced memory load and computational costs during the online test. Jie Qin 0004, Li Liu 0004, Ling Shao 0001, Bingbing Ni, Chen Chen 0001, Fumin Shen, Yunhong Wang 0001 |
CVPR | 2 |
| 2017 | Zero-Shot Action Recognition with Error-Correcting Output CodesabstractRecently, zero-shot action recognition (ZSAR) has emerged with the explosive growth of action categories. In this paper, we explore ZSAR from a novel perspective by adopting the Error-Correcting Output Codes (dubbed ZSECOC). Our ZSECOC equips the conventional ECOC with the additional capability of ZSAR, by addressing the domain shift problem. In particular, we learn discriminative ZSECOC for seen categories from both category-level semantics and intrinsic data structures. This procedure deals with domain shift implicitly by transferring the well-established correlations among seen categories to unseen ones. Moreover, a simple semantic transfer strategy is developed for explicitly transforming the learned embeddings of seen categories to better fit the underlying structure of unseen categories. As a consequence, our ZSECOC inherits the promising characteristics from ECOC as well as overcomes domain shift, making it more discriminative for ZSAR. We systematically evaluate ZSECOC on three realistic action benchmarks, i.e. Olympic Sports, HMDB51 and UCF101. The experimental results clearly show the superiority of ZSECOC over the state-of-the-art methods. Jie Qin 0004, Li Liu 0004, Ling Shao 0001, Fumin Shen, Bingbing Ni, Jiaxin Chen 0002, Yunhong Wang 0001 |
CVPR | 2 |
| 2017 | Deep Binaries: Encoding Semantic-Rich Cues for Efficient Textual-Visual Cross Retrieval
Yuming Shen, Li Liu 0004, Ling Shao 0001, Jingkuan Song |
ICCV | 2 |
| 2017 | DAP3D-Net: Where, what and how actions occur in videos?abstractAction parsing in videos with complex scenes is an interesting but challenging task in computer vision. In this paper, we propose a novel deep model based on 3D CNN (convolutional neural network) and LSTM (long short-term memory) module with a multi-task learning manner for effective Deep Action Parsing (DAP3D-Net) in videos. Particularly in the training phase, each action clip, sliced to several short consecutive segments, is fed into 3D CNN followed by LSTM to model the whole action dynamic information, so that action localization, classification and attributes learning can be jointly optimized via our deep model. Once the DAP3D-Net is trained, for an upcoming test video, we can describe each individual action in the video simultaneously as: Where the action occurs; What the action is and How the action is performed. To well demonstrate the effectiveness of the proposed DAP3D-Net, we also contribute a new Numerous-category Aligned Synthetic Action dataset, i.e., NASA, which consists of 200,000 action clips of 300 categories and with 33 pre-defined action attributes in two hierarchical levels (i.e., low-level attributes of basic body part movements and high-level attributes related to action motion). We learn DAP3D-Net using the NASA dataset and then evaluate it on our collected Human Action Understanding (HAU) dataset and the public THUMOS dataset. Experimental results show that our approach can accurately localize, categorize and describe multiple actions in realistic videos. Li Liu 0004, Yi Zhou 0007, Ling Shao 0001 |
ICRA | 1 |
| 2017 | Unsupervised Deep Video Hashing with Balanced RotationabstractRecently, hashing video contents for fast retrieval has received increasing attention due to the enormous growth of online videos. As the extension of image hashing techniques, traditional video hashing methods mainly focus on seeking the appropriate video features but pay little attention to how the video-specific features can be leveraged to achieve optimal binarization. In this paper, an end-to-end hashing framework, namely Unsupervised Deep Video Hashing (UDVH), is proposed, where feature extraction, balanced code learning and hash function learning are integrated and optimized in a self-taught manner. Particularly, distinguished from previous work, our framework enjoys two novelties: 1) an unsupervised hashing method that integrates the feature clustering and feature binarization, enabling the neighborhood structure to be preserved in the binary space; 2) a smart rotation applied to the video-specific features that are widely spread in the low-dimensional space such that the variance of dimensions can be balanced, thus generating more effective hash codes. Extensive experiments have been performed on two real-world datasets and the results demonstrate its superiority, compared to the state-of-the-art video hashing methods. To bootstrap further developments, the source code will be made publically available. Gengshen Wu, Li Liu 0004, Guiguang Ding, Jungong Han, Jialie Shen 0001, Ling Shao 0001 |
IJCAI | 2 |
| 2017 | Deep Asymmetric Pairwise HashingabstractRecently, deep neural networks based hashing methods have greatly improved the multimedia retrieval performance by simultaneously learning feature representations and binary hash functions. Inspired by the latest advance in the asymmetric hashing scheme, in this work, we propose a novel Deep Asymmetric Pairwise Hashing approach (DAPH) for supervised hashing. The core idea is that two deep convolutional models are jointly trained such that their output codes for a pair of images can well reveal the similarity indicated by their semantic labels. A pairwise loss is elaborately designed to preserve the pairwise similarities between images as well as incorporating the independence and balance hash code learning criteria. By taking advantage of the flexibility of asymmetric hash functions, we devise an efficient alternating algorithm to optimize the asymmetric deep hash functions and high-quality binary code jointly. Experiments on three image benchmarks show that DAPH achieves the state-of-the-art performance on large-scale image retrieval. Fumin Shen, Xin Gao 0001, Li Liu 0004, Yang Yang 0002, Heng Tao Shen |
ACM Multimedia | 3 |
| 2017 | Temporal Binary Coding for Large-Scale Video SearchabstractRecent years have witnessed the success of the emerging hash-based approximate nearest neighbor search techniques in large-scale image retrieval. However, for large-scale video search, most of the existing hashing methods mainly focus on the visual content contained in the still frames, without considering their temporal relations. Therefore, they usually suffer greatly from the insufficient capability of capturing the intrinsic video similarities, from both the visual and the temporal aspects. To address the problem, we propose a temporal binary coding solution in an unsupervised manner, which simultaneously considers the intrinsic relations among the visual content and the temporal consistency among the successive frames. To capture the inherent data similarities among videos, we adopt the sparse, nonnegative feature to characterize the common local visual content and approximate their intrinsic similarities using a low-rank matrix. Then a standard graph-based loss is adopted to guarantee that the learnt hash codes can well preserve the similarities. Furthermore, we introduce a subspace rotation to model the small variation among the successive frames, and thus essentially preserve the temporal consistency in Hamming space. Finally, we formulate the video hashing problem as a joint learning of the binary codes, the hash functions and the temporal variation, and devise an alternating optimization algorithm that enjoys fast training and discriminative hash functions. Extensive experiments on three large video datasets demonstrate the proposed method significantly outperforms a number of state-of-the-art hashing methods. Ke Xia, Yuqing Ma, Xianglong Liu 0001, Yadong Mu, Li Liu 0004 |
ACM Multimedia | 5 |
| 2017 | Classification by Retrieval: Binarizing Data and ClassifiersabstractThis paper proposes a generic formulation that significantly expedites the training and deployment of image classification models, particularly under the scenarios of many image categories and high feature dimensions. As the core idea, our method represents both the images and learned classifiers using binary hash codes, which are simultaneously learned from the training data. Classifying an image thereby reduces to retrieving its nearest class codes in the Hamming space. Specifically, we formulate multiclass image classification as an optimization problem over binary variables. The optimization alternatingly proceeds over the binary classifiers and image hash codes. Profiting from the special property of binary codes, we show that the sub-problems can be efficiently solved through either a binary quadratic program (BQP) or a linear program. In particular, for attacking the BQP problem, we propose a novel bit-flipping procedure which enjoys high efficacy and a local optimality guarantee. Our formulation supports a large family of empirical loss functions and is, in specific, instantiated by exponential and linear losses. Comprehensive evaluations are conducted on several representative image benchmarks. The experiments consistently exhibit reduced computational and memory complexities of model training and deployment, without sacrificing classification accuracy. Fumin Shen, Yadong Mu, Yang Yang 0002, Wei Liu 0005, Li Liu 0004, Jingkuan Song, Heng Tao Shen |
SIGIR | 5 |
| 2017 | Towards Fine-Grained Open Zero-Shot Learning: Inferring Unseen Visual Features from AttributesabstractZero-shot Learning (ZSL) can leverage attributes to recognise unseen instances. However, the training data is limited and cannot adequately discriminate fine-grained classes with similar attributes. In this paper, we propose a complementary procedure that inversely makes use of attributes to infer discriminative visual features for unseen classes. In this way, ZSL is fully converted into conventional supervised classification, where robust classifiers can be employed to address the fine-grained problem. To infer high-quality unseen data, we propose a novel algorithm named Orthogonal Semantic-Visual Embedding (OSVE) that can discover the tiny visual differences between different instances under the same attribute by an orthogonal embedding space. On two fine-grained benchmarks, CUB and SUN, our method remarkably improves the state-of-the-art results under standard ZSL settings. We further challenge the Open ZSL problem where the number of seen classes is significantly smaller than that of unseen classes. Substantial experiments manifest that the inferred visual features can be successfully fed to SVM which can effectively discriminate unseen classes from fine-grained open candidates. Yang Long 0001, Li Liu 0004, Ling Shao 0001 |
WACV | 2 |
| 2017 | Fast action retrieval from videos via feature disaggregation
Jie Qin 0004, Li Liu 0004, Mengyang Yu, Yunhong Wang 0001, Ling Shao 0001 |
Comput. Vis. Image Underst. | 2 |
| 2017 | Latent Structure Preserving HashingabstractAiming at efficient similarity search, hash functions are designed to embed high-dimensional feature descriptors to low-dimensional binary codes such that similar descriptors will lead to binary codes with a short distance in the Hamming space. It is critical to effectively maintain the intrinsic structure and preserve the original information of data in a hashing algorithm. In this paper, we propose a novel hashing algorithm called Latent Structure Preserving Hashing (LSPH), with the target of finding a well-structured low-dimensional data representation from the original high-dimensional data through a novel objective function based on Nonnegative Matrix Factorization (NMF) with their corresponding Kullback-Leibler divergence of data distribution as the regularization term. Via exploiting the joint probabilistic distribution of data, LSPH can automatically learn the latent information and successfully preserve the structure of high-dimensional data. To further achieve robust performance with complex and nonlinear data, in this paper, we also contribute a more generalized multi-layer LSPH (ML-LSPH) framework, in which hierarchical representations can be effectively learned by a multiplicative up-propagation algorithm. Once obtaining the latent representations, the hash functions can be easily acquired through multi-variable logistic regression. Experimental results on three large-scale retrieval datasets, i.e., SIFT 1M, GIST 1M and 500 K TinyImage, show that ML-LSPH can achieve better performance than the single-layer LSPH and both of them outperform existing hashing techniques on large-scale data. Li Liu 0004, Mengyang Yu, Ling Shao 0001 |
Int. J. Comput. Vis. | 1 |
| 2017 | Performance evaluation of deep feature learning for RGB-D image/video classification
Ling Shao 0001, Ziyun Cai, Li Liu 0004, Ke Lu 0002 |
Inf. Sci. | 3 |
| 2017 | RGB-D datasets using microsoft kinect or similar sensors: a surveyabstractRGB-D data has turned out to be a very useful representation of an indoor scene for solving fundamental computer vision problems. It takes the advantages of the color image that provides appearance information of an object and also the depth image that is immune to the variations in color, illumination, rotation angle and scale. With the invention of the low-cost Microsoft Kinect sensor, which was initially used for gaming and later became a popular device for computer vision, high quality RGB-D data can be acquired easily. In recent years, more and more RGB-D image/video datasets dedicated to various applications have become available, which are of great importance to benchmark the state-of-the-art. In this paper, we systematically survey popular RGB-D datasets for different applications including object recognition, scene classification, hand gesture recognition, 3D-simultaneous localization and mapping, and pose estimation. We provide the insights into the characteristics of each important dataset, and compare the popularity and the difficulty of those datasets. Overall, the main goal of this survey is to give a comprehensive description about the available RGB-D datasets and thus to guide researchers in the selection of suitable datasets for evaluating their algorithms. Ziyun Cai, Jungong Han, Li Liu 0004, Ling Shao 0001 |
Multim. Tools Appl. | 3 |
| 2017 | Learning to Hash With Optimized Anchor Embedding for Scalable RetrievalabstractSparse representation and image hashing are powerful tools for data representation and image retrieval respectively. The combinations of these two tools for scalable image retrieval, i.e., sparse hashing (SH) methods, have been proposed in recent years and the preliminary results are promising. The core of those methods is a scheme that can efficiently embed the (high-dimensional) image features into a low-dimensional Hamming space, while preserving the similarity between features. Existing SH methods mostly focus on finding better sparse representations of images in the hash space. We argue that the anchor set utilized in sparse representation is also crucial, which was unfortunately underestimated by the prior art. To this end, we propose a novel SH method that optimizes the integration of the anchors, such that the features can be better embedded and binarized, termed as Sparse Hashing with Optimized Anchor Embedding. The central idea is to push the anchors far from the axis while preserving their relative positions so as to generate similar hashcodes for neighboring features. We formulate this idea as an orthogonality constrained maximization problem and an efficient and novel optimization framework is systematically exploited. Extensive experiments on five benchmark image data sets demonstrate that our method outperforms several state-of-the-art related methods. Guiguang Ding, Li Liu 0004, Jungong Han, Ling Shao 0001 |
IEEE Trans. Image Process. | 3 |
| 2017 | Sequential Discrete Hashing for Scalable Cross-Modality Similarity RetrievalabstractWith the dramatic development of the Internet, how to exploit large-scale retrieval techniques for multimodal web data has become one of the most popular but challenging problems in computer vision and multimedia. Recently, hashing methods are used for fast nearest neighbor search in large-scale data spaces, by embedding high-dimensional feature descriptors into a similarity preserving Hamming space with a low dimension. Inspired by this, in this paper, we introduce a novel supervised cross-modality hashing framework, which can generate unified binary codes for instances represented in different modalities. Particularly, in the learning phase, each bit of a code can be sequentially learned with a discrete optimization scheme that jointly minimizes its empirical loss based on a boosting strategy. In a bitwise manner, hash functions are then learned for each modality, mapping the corresponding representations into unified hash codes. We regard this approach as cross-modality sequential discrete hashing (CSDH), which can effectively reduce the quantization errors arisen in the oversimplified rounding-off step and thus lead to high-quality binary codes. In the test phase, a simple fusion scheme is utilized to generate a unified hash code for final retrieval by merging the predicted hashing results of an unseen instance from different modalities. The proposed CSDH has been systematically evaluated on three standard data sets: Wiki, MIRFlickr, and NUS-WIDE, and the results show that our method significantly outperforms the state-of-the-art multimodality hashing techniques. Li Liu 0004, Zijia Lin, Ling Shao 0001, Fumin Shen, Guiguang Ding, Jungong Han |
IEEE Trans. Image Process. | 1 |
| 2017 | Learning Short Binary Codes for Large-scale Image RetrievalabstractLarge-scale visual information retrieval has become an active research area in this big data era. Recently, hashing/binary coding algorithms prove to be effective for scalable retrieval applications. Most existing hashing methods require relatively long binary codes (i.e., over hundreds of bits, sometimes even thousands of bits) to achieve reasonable retrieval accuracies. However, for some realistic and unique applications, such as on wearable or mobile devices, only short binary codes can be used for efficient image retrieval due to the limitation of computational resources or bandwidth on these devices. In this paper, we propose a novel unsupervised hashing approach called min-cost ranking (MCR) specifically for learning powerful short binary codes (i.e., usually the code length shorter than 100 b) for scalable image retrieval tasks. By exploring the discriminative ability of each dimension of data, MCR can generate one bit binary code for each dimension and simultaneously rank the discriminative separability of each bit according to the proposed cost function. Only top-ranked bits with minimum cost-values are then selected and grouped together to compose the final salient binary codes. Extensive experimental results on large-scale retrieval demonstrate that MCR can achieve comparative performance as the state-of-the-art hashing algorithms but with significantly shorter codes, leading to much faster large-scale retrieval. Li Liu 0004, Mengyang Yu, Ling Shao 0001 |
IEEE Trans. Image Process. | 1 |
| 2017 | Asymmetric Binary Coding for Image SearchabstractLearning to hash has attracted broad research interests in recent computer vision and machine learning studies, due to its ability to accomplish efficient approximate nearest neighbor search. However, the closely related task, maximum inner product search (MIPS), has rarely been studied in this literature. To facilitate the MIPS study, in this paper, we introduce a general binary coding framework based on asymmetric hash functions, named asymmetric inner-product binary coding (AIBC). In particular, AIBC learns two different hash functions, which can reveal the inner products between original data vectors by the generated binary vectors. Although conceptually simple, the associated optimization is very challenging due to the highly nonsmooth nature of the objective that involves sign functions. We tackle the nonsmooth optimization in an alternating manner, by which each single coding function is optimized in an efficient discrete manner. We also simplify the objective by discarding the quadratic regularization term which significantly boosts the learning efficiency. Both problems are optimized in an effective discrete way without continuous relaxations, which produces high-quality hash codes. In addition, we extend the AIBC approach to the supervised hashing scenario, where the inner products of learned binary codes are forced to fit the supervised similarities. Extensive experiments on several benchmark image retrieval databases validate the superiority of the AIBC approaches over many recently proposed hashing algorithms. Fumin Shen, Yang Yang 0002, Li Liu 0004, Wei Liu 0005, Dacheng Tao, Heng Tao Shen |
IEEE Trans. Multim. | 3 |
| 2017 | Binary Set Embedding for Cross-Modal RetrievalabstractCross-modal retrieval is such a challenging topic that traditional global representations would fail to bridge the semantic gap between images and texts to a satisfactory level. Using local features from images and words from documents directly can be more robust for the scenario with large intraclass variations and small interclass discrepancies. In this paper, we propose a novel unsupervised binary coding algorithm called binary set embedding (BSE) to obtain meaningful hash codes for local features from the image domain and words from text domain. Understanding image features with the word vectors learned from the human language instead of the provided documents from data sets, BSE can map samples into a common Hamming space effectively and efficiently where each sample is represented by the sets of local feature descriptors from image and text domains. In particular, BSE explores relationship among local features in both feature level and image (text) level, which can balance the sensitivity of each other. Furthermore, a recursive orthogonalization procedure is applied to reduce the redundancy of codes. Extensive experiments demonstrate the superior performance of BSE compared with state-of-the-art cross-modal hashing methods using either image or text queries. Mengyang Yu, Li Liu 0004, Ling Shao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2016 | Attribute Embedding with Visual-Semantic Ambiguity Removal for Zero-shot Learning
Yang Long 0001, Li Liu 0004, Ling Shao 0001 |
BMVC | 2 |
| 2016 | DAVE: A Unified Framework for Fast Vehicle Detection and Annotation
Yi Zhou 0007, Li Liu 0004, Ling Shao 0001, Matt Mellor |
ECCV (2) | 2 |
| 2016 | Kernelized Multiview Projection for Robust Action RecognitionabstractConventional action recognition algorithms adopt a single type of feature or a simple concatenation of multiple features. In this paper, we propose to better fuse and embed different feature representations for action recognition using a novel spectral coding algorithm called Kernelized Multiview Projection (KMP). Computing the kernel matrices from different features/views via time-sequential distance learning, KMP can encode different features with different weights to achieve a low-dimensional and semantically meaningful subspace where the distribution of each view is sufficiently smooth and discriminative. More crucially, KMP is linear for the reproducing kernel Hilbert space, which allows it to be competent for various practical applications. We demonstrate KMP’s performance for action recognition on five popular action datasets and the results are consistently superior to state-of-the-art techniques. Ling Shao 0001, Li Liu 0004, Mengyang Yu |
Int. J. Comput. Vis. | 2 |
| 2016 | Structure-Preserving Binary Representations for RGB-D Action RecognitionabstractIn this paper, we propose a novel binary local representation for RGB-D video data fusion with a structure-preserving projection. Our contribution consists of two aspects. Toacquire a general feature for the video data, we convert the problem to describing the gradient fields of RGB and depth information of video sequences. With the local fluxes of the gradient fields, which include the orientation and the magnitude of the neighborhood of each point, a new kind of continuous local descriptor called Local Flux Feature(LFF) is obtained. Then the LFFs from RGB and depth channels are fused into a Hamming space via the Structure Preserving Projection (SPP). Specifically, an orthogonal projection matrix is applied to preserve the pairwise structure with a shape constraint to avoid the collapse of data structure in the projected space. Furthermore, a bipartite graph structure of data is taken into consideration, which is regarded as a higher level connection between samples and classes than the pairwise structure of local features. Theextensive experiments show not only the high efficiency of binary codes and the effectiveness of combining LFFs from RGB-D channels via SPP on various action recognition benchmarks of RGB-D data, but also the potential power of LFF for general action recognition. Mengyang Yu, Li Liu 0004, Ling Shao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2016 | Beyond Semantic Attributes: Discrete Latent Attributes Learning for Zero-Shot RecognitionabstractIn this letter, we propose a novel approach for learning semantics-driven attributes, which are discriminative for zero-shot visual recognition. Latent attributes are derived in a principled manner, aiming at maintaining class-level semantic relatedness and attribute-wise balancedness. Unlike existing methods that binarize learned real-valued attributes via a quantization stage, we directly learn the optimal binary attributes by effectively addressing a discrete optimization problem. Particularly, we propose a class-wise discrete descent algorithm, based on which latent attributes of each class are learned iteratively. Moreover, we propose to simultaneously predict multiple attributes from low-level features via multioutput neural networks (MONN), which can model intrinsic correlation among attributes and make prediction more tractable. Extensive experiments on two standard datasets clearly demonstrate the superiority of our method over the state-of-the-arts. Jie Qin 0004, Yunhong Wang 0001, Li Liu 0004, Jiaxin Chen 0002, Ling Shao 0001 |
IEEE Signal Process. Lett. | 3 |
| 2016 | Unsupervised Local Feature Hashing for Image Similarity SearchabstractThe potential value of hashing techniques has led to it becoming one of the most active research areas in computer vision and multimedia. However, most existing hashing methods for image search and retrieval are based on global feature representations, which are susceptible to image variations such as viewpoint changes and background cluttering. Traditional global representations gather local features directly to output a single vector without the analysis of the intrinsic geometric property of local features. In this paper, we propose a novel unsupervised hashing method called unsupervised bilinear local hashing (UBLH) for projecting local feature descriptors from a high-dimensional feature space to a lower-dimensional Hamming space via compact bilinear projections rather than a single large projection matrix. UBLH takes the matrix expression of local features as input and preserves the feature-to-feature and image-to-image structures of local features simultaneously. Experimental results on challenging data sets including Caltech-256, SUN397, and Flickr 1M demonstrate the superiority of UBLH compared with state-of-the-art hashing methods. Li Liu 0004, Mengyang Yu, Ling Shao 0001 |
IEEE Trans. Cybern. | 1 |
| 2016 | Learning Spatio-Temporal Representations for Action Recognition: A Genetic Programming ApproachabstractExtracting discriminative and robust features from video sequences is the first and most critical step in human action recognition. In this paper, instead of using handcrafted features, we automatically learn spatio-temporal motion features for action recognition. This is achieved via an evolutionary method, i.e., genetic programming (GP), which evolves the motion feature descriptor on a population of primitive 3D operators (e.g., 3D-Gabor and wavelet). In this way, the scale and shift invariant features can be effectively extracted from both color and optical flow sequences. We intend to learn data adaptive descriptors for different datasets with multiple layers, which makes fully use of the knowledge to mimic the physical structure of the human visual cortex for action recognition and simultaneously reduce the GP searching space to effectively accelerate the convergence of optimal solutions. In our evolutionary architecture, the average cross-validation classification error, which is calculated by an support-vector-machine classifier on the training set, is adopted as the evaluation criterion for the GP fitness function. After the entire evolution procedure finishes, the best-so-far solution selected by GP is regarded as the (near-)optimal action descriptor obtained. The GP-evolving feature extraction method is evaluated on four popular action datasets, namely KTH, HMDB51, UCF YouTube, and Hollywood2. Experimental results show that our method significantly outperforms other types of features, either hand-designed or machine-learned. Li Liu 0004, Ling Shao 0001, Xuelong Li 0001, Ke Lu 0002 |
IEEE Trans. Cybern. | 1 |
| 2016 | Compressive Sequential Learning for Action Similarity LabelingabstractHuman action recognition in videos has been extensively studied in recent years due to its wide range of applications. Instead of classifying video sequences into a number of action categories, in this paper, we focus on a particular problem of action similarity labeling (ASLAN), which aims at verifying whether a pair of videos contain the same type of action or not. To address this challenge, a novel approach called compressive sequential learning (CSL) is proposed by leveraging the compressive sensing theory and sequential learning. We first project data points to a low-dimensional space by effectively exploring an important property in compressive sensing: the restricted isometry property. In particular, a very sparse measurement matrix is adopted to reduce the dimensionality efficiently. We then learn an ensemble classifier for measuring similarities between pairwise videos by iteratively minimizing its empirical risk with the AdaBoost strategy on the training set. Unlike conventional AdaBoost, the weak learner for each iteration is not explicitly defined and its parameters are learned through greedy optimization. Furthermore, an alternative of CSL named compressive sequential encoding is developed as an encoding technique and followed by a linear classifier to address the similarity-labeling problem. Our method has been systematically evaluated on four action data sets: ASLAN, KTH, HMDB51, and Hollywood2, and the results show the effectiveness and superiority of our method for ASLAN. Jie Qin 0004, Li Liu 0004, Zhaoxiang Zhang 0001, Yunhong Wang 0001, Ling Shao 0001 |
IEEE Trans. Image Process. | 2 |
| 2016 | Sequential Compact Code Learning for Unsupervised Image HashingabstractEffective hashing for large-scale image databases is a popular research area, attracting much attention in computer vision and visual information retrieval. Several recent methods attempt to learn either graph embedding or semantic coding for fast and accurate applications. In this paper, a novel unsupervised framework, termed evolutionary compact embedding (ECE), is introduced to automatically learn the task-specific binary hash codes. It can be regarded as an optimization algorithm that combines the genetic programming (GP) and a boosting trick. In our architecture, each bit of ECE is iteratively computed using a weak binary classification function, which is generated through GP evolving by jointly minimizing its empirical risk with the AdaBoost strategy on a training set. We address this as greedy optimization by embedding high-dimensional data points into a similarity-preserved Hamming space with a low dimension. We systematically evaluate ECE on two data sets, SIFT 1M and GIST 1M, showing the effectiveness and the accuracy of our method for a large-scale similarity search. Li Liu 0004, Ling Shao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2015 | Local Feature Binary Coding for Approximate Nearest Neighbor SearchabstractThe potential value of hashing techniques has led to it becoming one of the most active research areas in computer vision and multimedia. However, most existing hashing methods for image search and retrieval are based on global representations, e.g., GIST, which lack the analysis of the intrinsic geometric property of local features and heavily limit the effectiveness of the hash code. In this paper, we propose a novel supervised hashing method called Local Feature Binary Coding (LFBC) for projecting local feature descriptors from a high-dimensional feature space to a lower-dimensional Hamming space via compact bilinear projections rather than a single large projection matrix. LFBC takes the matrix expression of local features as input and preserves the feature-to-feature and image-to-class structures simultaneously. Experimental results on challenging datasets including Caltech-256, SUN397 and NUS-WIDE demonstrate the superiority of LFBC compared with state-of-the-art hashing methods. Li Liu 0004, Mengyang Yu, Ling Shao 0001 |
BMVC | 1 |
| 2015 | Latent Structure Preserving Hashing
Ziyun Cai, Li Liu 0004, Mengyang Yu, Ling Shao 0001 |
BMVC | 2 |
| 2015 | Fast Action Retrieval from Videos via Feature DisaggregationabstractLearning based hashing methods, which aim at learning similarity-preserving binary codes for efficient nearest neighbor search, have been actively studied recently. A majority of the approaches address hashing problems for image collections. However, due to the extra temporal information, videos are usually represented by much higher dimensional (thousands or even more) features compared with images, causing high computational complexity for conventional hashing schemes. In this paper, we propose a simple and efficient hashing scheme for high-dimensional video data. This method, called Disaggregation Hashing, exploits the correlations among different feature dimensions. An intuitive feature disaggregation method is first proposed, followed by a novel hashing algorithm based on different feature clusters. We demonstrate the efficiency and effectiveness of our method by theoretical analysis and exploring its application on action retrieval from video databases. Extensive experiments show the superiority of our binary coding scheme over state-of-the-art hashing methods. Jie Qin 0004, Li Liu 0004, Mengyang Yu, Yunhong Wang 0001, Ling Shao 0001 |
BMVC | 2 |
| 2015 | Projection Bank: From High-Dimensional Data to Medium-Length Binary CodesabstractRecently, very high-dimensional feature representations, e.g., Fisher Vector, have achieved excellent performance for visual recognition and retrieval. However, these lengthy representations always cause extremely heavy computational and storage costs and even become unfeasible in some large-scale applications. A few existing techniques can transfer very high-dimensional data into binary codes, but they still require the reduced code length to be relatively long to maintain acceptable accuracies. To target a better balance between computational efficiency and accuracies, in this paper, we propose a novel embedding method called Binary Projection Bank (BPB), which can effectively reduce the very high-dimensional representations to medium-dimensional binary codes without sacrificing accuracies. Instead of using conventional single linear or bilinear projections, the proposed method learns a bank of small projections via the max-margin constraint to optimally preserve the intrinsic data similarity. We have systematically evaluated the proposed method on three datasets: Flickr 1M, ILSVR2010 and UCF101, showing competitive retrieval and recognition accuracies compared with state-of-the-art approaches, but with a significantly smaller memory footprint and lower coding complexity. Li Liu 0004, Mengyang Yu, Ling Shao 0001 |
ICCV | 1 |
| 2015 | Image-to-class distance ratio: A feature filtering metric for image classification
Shoubiao Tan, Li Liu 0004, Chunyu Peng, Ling Shao 0001 |
Neurocomputing | 2 |
| 2015 | Evolutionary compact embedding for large-scale image classification
Li Liu 0004, Ling Shao 0001, Xuelong Li 0001 |
Inf. Sci. | 1 |
| 2015 | Multiview Alignment Hashing for Efficient Image SearchabstractHashing is a popular and efficient method for nearest neighbor search in large-scale data spaces by embedding high-dimensional feature descriptors into a similarity preserving Hamming space with a low dimension. For most hashing methods, the performance of retrieval heavily depends on the choice of the high-dimensional feature descriptor. Furthermore, a single type of feature cannot be descriptive enough for different images when it is used for hashing. Thus, how to combine multiple representations for learning effective hashing functions is an imminent task. In this paper, we present a novel unsupervised multiview alignment hashing approach based on regularized kernel nonnegative matrix factorization, which can find a compact representation uncovering the hidden semantics and simultaneously respecting the joint probability distribution of data. In particular, we aim to seek a matrix factorization to effectively fuse the multiple information sources meanwhile discarding the feature redundancy. Since the raised problem is regarded as nonconvex and discrete, our objective function is then optimized via an alternate way with relaxation and converges to a locally optimal solution. After finding the low-dimensional representation, the hashing functions are finally obtained through multivariable logistic regression. The proposed method is systematically evaluated on three data sets: 1) Caltech-256; 2) CIFAR-10; and 3) CIFAR-20, and the results show that our method significantly outperforms the state-of-the-art multiview hashing techniques. Li Liu 0004, Mengyang Yu, Ling Shao 0001 |
IEEE Trans. Image Process. | 1 |
| 2014 | Discriminative Partition Sparsity AnalysisabstractEffective dimensionality reduction has been an attractive research area for many large-scale vision and multimedia tasks. Several recent methods attempt to learn optimized graph-based embedding for fast and accurate applications. In this paper, we propose a novel linear unsupervised algorithm, termed Discriminative Partition Sparsity Analysis (DPSA), explicitly considering different probabilistic distributions that exist over the data points, meanwhile preserving the natural locality relationship among the data. Specifically, the Gaussian mixture model (GMM) is first applied to partition all samples into several clusters. In each cluster, a number of sparse sub-graphs are computed via the ℓ1-norm constraint to optimally represent the intrinsic data structure. Such sub-graphs are demonstrated to be robust to data noise, automatically sparse and adaptive to the neighborhood. All the sub-graphs from the clusters are then combined into a whole discriminative optimization framework for final reduction. We have systematically evaluated our method on three image datasets: USPS digital hand-writing, CMU PIE face and CIFAR-10 tiny image, showing its accurate and robust performance for image classification. Li Liu 0004, Ling Shao 0001 |
ICPR | 1 |
| 2014 | Realistic action recognition via sparsely-constructed Gaussian processes
Li Liu 0004, Ling Shao 0001, Feng Zheng 0001, Xuelong Li 0001 |
Pattern Recognit. | 1 |
| 2014 | Natural image denoising using evolved local adaptive filters
Ruomei Yan, Ling Shao 0001, Li Liu 0004, Yan Liu 0004 |
Signal Process. | 3 |
| 2014 | Feature Learning for Image Classification Via Multiobjective Genetic ProgrammingabstractFeature extraction is the first and most critical step in image classification. Most existing image classification methods use hand-crafted features, which are not adaptive for different image domains. In this paper, we develop an evolutionary learning methodology to automatically generate domain-adaptive global feature descriptors for image classification using multiobjective genetic programming (MOGP). In our architecture, a set of primitive 2-D operators are randomly combined to construct feature descriptors through the MOGP evolving and then evaluated by two objective fitness criteria, i.e., the classification error and the tree complexity. After the entire evolution procedure finishes, the best-so-far solution selected by the MOGP is regarded as the (near-)optimal feature descriptor obtained. To evaluate its performance, the proposed approach is systematically tested on the Caltech-101, the MIT urban and nature scene, the CMU PIE, and Jochen Triesch Static Hand Posture II data sets, respectively. Experimental results verify that our method significantly outperforms many state-of-the-art hand-designed features and two feature learning techniques in terms of classification accuracy. Ling Shao 0001, Li Liu 0004, Xuelong Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2013 | Learning Discriminative Representations from RGB-D Video Data
Li Liu 0004, Ling Shao 0001 |
IJCAI | 1 |
| 2013 | Building holistic descriptors for scene recognition: a multi-objective genetic programming approachabstractReal-world scene recognition has been one of the most challenging research topics in computer vision, due to the tremendous intraclass variability and the wide range of scene categories. In this paper, we successfully apply an evolutionary methodology to automatically synthesize domain-adaptive holistic descriptors for the task of scene recognition, instead of using hand-tuned descriptors. We address this as an optimization problem by using multi-objective genetic programming (MOGP). Specifically, a set of primitive operators and filters are first randomly assembled in theMOGP framework as tree-based combinations, which are then evaluated by two objective fitness criteria i.e., the classification error and the tree complexity. Finally, the best-so-far solution selected by MOGP is regarded as the (near-)optimal feature descriptor for scene recognition. We have evaluated our approach on three realistic scene datasets: MIT urban and nature, SUN and UIUC Sport. Experimental results consistently show that our MOGP-generated descriptors achieve significantly higher recognition accuracies compared with state-of-the-art hand-crafted and machine-learned features. Li Liu 0004, Ling Shao 0001, Xuelong Li 0001 |
ACM Multimedia | 1 |
| 2013 | Boosted key-frame selection and correlated pyramidal motion-feature representation for human action recognition
Li Liu 0004, Ling Shao 0001, Peter I. Rockett |
Pattern Recognit. | 1 |
| 2013 | Human action recognition based on boosted feature selection and naive Bayes nearest-neighbor classification
Li Liu 0004, Ling Shao 0001, Peter I. Rockett |
Signal Process. | 1 |
| 2013 | Learning Discriminative Key Poses for Action RecognitionabstractIn this paper, we present a new approach for human action recognition based on key-pose selection and representation. Poses in video frames are described by the proposed extensive pyramidal features (EPFs), which include the Gabor, Gaussian, and wavelet pyramids. These features are able to encode the orientation, intensity, and contour information and therefore provide an informative representation of human poses. Due to the fact that not all poses in a sequence are discriminative and representative, we further utilize the AdaBoost algorithm to learn a subset of discriminative poses. Given the boosted poses for each video sequence, a new classifier named weighted local naive Bayes nearest neighbor is proposed for the final action classification, which is demonstrated to be more accurate and robust than other classifiers, e.g., support vector machine (SVM) and naive Bayes nearest neighbor. The proposed method is systematically evaluated on the KTH data set, the Weizmann data set, the multiview IXMAS data set, and the challenging HMDB51 data set. Experimental results manifest that our method outperforms the state-of-the-art techniques in terms of recognition rate. Li Liu 0004, Ling Shao 0001, Xiantong Zhen, Xuelong Li 0001 |
IEEE Trans. Cybern. | 1 |
| 2012 | Genetic Programming-Evolved Spatio-Temporal Descriptor for Human Action RecognitionabstractThe potential value of human action recognition has led to it becoming one of the most active research subjects in computer vision. In this paper, we propose a novel method to automatically generate low-level spatio-temporal descriptors showing good performance, for high-level human-action recognition tasks. We address this as an optimization problem using genetic programming (GP), an evolutionary method, which produces the descriptor by combining a set of primitive 3D operators. As far as we are aware, this is the first report of using GP for evolving spatio-temporal descriptors for action recognition. In our evolutionary architecture, the average cross-validation classification error calculated using the support-vector machine (SVM) classifier is used as the GP fitness function. We run GP on a mixed dataset combining the KTH and the Weizmann datasets to obtain a promising feature-descriptor solution for action recognition. To demonstrate generalizability, the best descriptor generated so far by GP has also been tested on the IXMAS dataset leading to better accuracies compared with some previous hand-crafted descriptors. Li Liu 0004, Ling Shao 0001, Peter I. Rockett |
BMVC | 1 |