VLDB 2026 Research / reviewers in the wild / expert
Fan Zhu 0001
dblp:96/2445-1
· DBLP profile ↗
76ranked-venue papers
11as first author
9since 2021 · last 2023
0000-0002-2009-1152ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 59 · 10 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 48 · 7 first-author · 3 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
38 papers |
Deep learning architectures and training · 24% Face, body and person analysis · 11% 3D vision · 10% | |
| Computer graphics and multimedia
13 papers |
Image and video processing · 45% Multimedia analysis and retrieval · 41% Geometric modeling and processing · 14% | |
| Databases, data mining, and information retrieval
9 papers |
Information retrieval · 44% Query processing and optimization · 18% Data mining · 14% |
Topics — the 30 heaviest of 113, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Multimedia analysis and retrieval
3d shape retrieval |
2.1 | 8 | 2019 | Deep Sketch-Shape Hashing With Segmented 3D Stochastic Viewing · CVPR 2019 DeepShape: Deep-Learned Shape Descriptor for 3D Shape Retrieval · IEEE Trans. Pattern Anal. Mach. Intell. 2017 Learning Barycentric Representations of 3D Shapes for Sketch-Based 3D Shape Retrieval · CVPR 2017 |
Computer vision › Face, body and person analysis
person re-identification |
1.7 | 4 | 2023 | Learning Multi-Attention Context Graph for Group-Based Re-Identification · IEEE Trans. Pattern Anal. Mach. Intell. 2023 Deep Local Binary Coding for Person Re-Identification by Delving into the Details · ACM Multimedia 2020 Learning Multi-Granular Hypergraphs for Video-Based Person Re-Identification · CVPR 2020 |
Machine learning › Deep learning architectures and training
normalization |
1.6 | 3 | 2023 | Normalization Techniques in Training DNNs: Methodology, Analysis and Application · IEEE Trans. Pattern Anal. Mach. Intell. 2023 Group Whitening: Balancing Learning Efficiency and Representational Capacity · CVPR 2021 An Investigation Into the Stochasticity of Batch Whitening · CVPR 2020 |
Machine learning › Transfer learning and domain adaptation
zero-shot learning |
1.2 | 3 | 2020 | Zero-VAE-GAN: Generating Unseen Features for Generalized and Transductive Zero-Shot Learning · IEEE Trans. Image Process. 2020 Region Graph Embedding Network for Zero-Shot Learning · ECCV (4) 2020 Attentive Region Embedding Network for Zero-Shot Learning · CVPR 2019 |
Multimedia analysis and retrieval › 3d shape retrieval
sketch-based 3d shape retrieval |
1.2 | 4 | 2019 | Deep Sketch-Shape Hashing With Segmented 3D Stochastic Viewing · CVPR 2019 Learning Barycentric Representations of 3D Shapes for Sketch-Based 3D Shape Retrieval · CVPR 2017 Deep Correlated Metric Learning for Sketch-based 3D Shape Retrieval · AAAI 2017 |
Machine learning › Deep learning architectures and training
attention mechanism |
0.9 | 3 | 2020 | Learning Attentive and Hierarchical Representations for 3D Shape Recognition · ECCV (15) 2020 Relational Attention Network for Crowd Counting · ICCV 2019 Attentional Neural Fields for Crowd Counting · ICCV 2019 |
Machine learning › Representation and self-supervised learning › redundancy reduction
whitening |
0.9 | 2 | 2021 | Group Whitening: Balancing Learning Efficiency and Representational Capacity · CVPR 2021 Iterative Normalization: Beyond Standardization Towards Efficient Whitening · CVPR 2019 |
Query processing and optimization › join processing
set containment join |
0.9 | 2 | 2021 | Internal and external memory set containment join · VLDB J. 2021 LCJoin: Set Containment Join via List Crosscutting · ICDE 2019 |
Image and video processing › image restoration
image denoising |
0.9 | 2 | 2020 | Noisy-as-Clean: Learning Self-Supervised Denoising From Corrupted Image · IEEE Trans. Image Process. 2020 NLH: A Blind Pixel-Level Non-Local Method for Real-World Image Denoising · IEEE Trans. Image Process. 2020 |
Geometric modeling and processing
shape matching |
0.8 | 3 | 2017 | DeepShape: Deep-Learned Shape Descriptor for 3D Shape Retrieval · IEEE Trans. Pattern Anal. Mach. Intell. 2017 Learning Barycentric Representations of 3D Shapes for Sketch-Based 3D Shape Retrieval · CVPR 2017 Deepshape: Deep learned shape descriptor for 3D shape matching and retrieval · CVPR 2015 |
Machine learning › Deep learning architectures and training
convolutional neural network |
0.8 | 2 | 2020 | On the Number of Linear Regions of Convolutional Neural Networks · ICML 2020 TBN: Convolutional Neural Network with Ternary Inputs and Binary Weights · ECCV (2) 2018 |
Computer vision › Image recognition and object detection › object counting
crowd counting |
0.8 | 2 | 2019 | Attentional Neural Fields for Crowd Counting · ICCV 2019 Relational Attention Network for Crowd Counting · ICCV 2019 |
Machine learning › Graph learning
graph neural network |
0.7 | 1 | 2023 | Learning Multi-Attention Context Graph for Group-Based Re-Identification · IEEE Trans. Pattern Anal. Mach. Intell. 2023 |
Computer vision › Face, body and person analysis › person re-identification
group re-identification |
0.7 | 1 | 2023 | Learning Multi-Attention Context Graph for Group-Based Re-Identification · IEEE Trans. Pattern Anal. Mach. Intell. 2023 |
Mathematical optimization › integer programming
binary optimization |
0.6 | 1 | 2022 | A Generalized Method for Binary Optimization: Convergence Analysis and Applications · IEEE Trans. Pattern Anal. Mach. Intell. 2022 |
Mathematical optimization
convergence analysis |
0.6 | 1 | 2022 | A Generalized Method for Binary Optimization: Convergence Analysis and Applications · IEEE Trans. Pattern Anal. Mach. Intell. 2022 |
Computer vision › Image recognition and object detection › object detection
anchor-free object detection |
0.5 | 1 | 2021 | Anchor-Free Person Search · CVPR 2021 |
Computer vision › Vision and language
cross-modal matching |
0.5 | 1 | 2021 | P2-Net: Joint Description and Detection of Local Features for Pixel and Point Matching · ICCV 2021 |
Computer vision › 3D vision › feature matching
local feature detection and description |
0.5 | 1 | 2021 | P2-Net: Joint Description and Detection of Local Features for Pixel and Point Matching · ICCV 2021 |
Computer vision › Face, body and person analysis
person search |
0.5 | 1 | 2021 | Anchor-Free Person Search · CVPR 2021 |
Query processing and optimization
join processing |
0.5 | 1 | 2021 | Internal and external memory set containment join · VLDB J. 2021 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
density estimation |
0.5 | 2 | 2019 | Attentional Neural Fields for Crowd Counting · ICCV 2019 Relational Attention Network for Crowd Counting · ICCV 2019 |
Machine learning › Transfer learning and domain adaptation
cross-domain learning |
0.4 | 2 | 2016 | Color object recognition via cross-domain learning on RGB-D images · ICRA 2016 Weakly-Supervised Cross-Domain Dictionary Learning for Visual Recognition · Int. J. Comput. Vis. 2014 |
Computer vision › 3D vision › 3d shape analysis
3d shape recognition |
0.4 | 1 | 2020 | Learning Attentive and Hierarchical Representations for 3D Shape Recognition · ECCV (15) 2020 |
Machine learning › Representation and self-supervised learning › representation learning › neural network representation learning › deep representation learning
autoencoder representation learning |
0.4 | 1 | 2020 | Auto-Encoding Twin-Bottleneck Hashing · CVPR 2020 |
Machine learning › Representation and self-supervised learning › representation learning › discrete representation learning
binary representation learning |
0.4 | 1 | 2020 | Deep Local Binary Coding for Person Re-Identification by Delving into the Details · ACM Multimedia 2020 |
Machine learning › Generative modeling
feature generation |
0.4 | 1 | 2020 | Zero-VAE-GAN: Generating Unseen Features for Generalized and Transductive Zero-Shot Learning · IEEE Trans. Image Process. 2020 |
Machine learning › Transfer learning and domain adaptation › zero-shot learning
generalized zero-shot learning |
0.4 | 1 | 2020 | Zero-VAE-GAN: Generating Unseen Features for Generalized and Transductive Zero-Shot Learning · IEEE Trans. Image Process. 2020 |
Machine learning › Graph learning
hypergraph learning |
0.4 | 1 | 2020 | Learning Multi-Granular Hypergraphs for Video-Based Person Re-Identification · CVPR 2020 |
Machine learning › Graph learning
network embedding |
0.4 | 1 | 2020 | Region Graph Embedding Network for Zero-Shot Learning · ECCV (4) 2020 |
Methods — techniques the papers use, named apart from their topics
multi-view learning · 1.5alternating optimization · 1.5attention mechanism · 1.1newton's iteration · 0.8multi-instance learning · 0.8hierarchical representation learning · 0.7self-attention · 0.7normalization representation recovery · 0.7normalization operation · 0.7normalization area partitioning · 0.7multi-level attention · 0.7matrix perturbation · 0.6multiscale shape distribution · 0.5fisher discrimination criterion · 0.5deep autoencoder · 0.5set containment join · 0.5external-memory algorithm · 0.5wiener filtering · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Normalization Techniques in Training DNNs: Methodology, Analysis and ApplicationabstractNormalization techniques are essential for accelerating the training and improving the generalization of deep neural networks (DNNs), and have successfully been used in various applications. This paper reviews and comments on the past, present and future of normalization methods in the context of DNN training. We provide a unified picture of the main motivation behind different approaches from the perspective of optimization, and present a taxonomy for understanding the similarities and differences between them. Specifically, we decompose the pipeline of the most representative normalizing activation methods into three components: the normalization area partitioning, normalization operation and normalization representation recovery. In doing so, we provide insight for designing new normalization technique. Finally, we discuss the current progress in understanding normalization methods, and provide a comprehensive review of the applications of normalization for particular tasks, in which it can effectively solve the key issues. Lei Huang 0015, Jie Qin 0004, Yi Zhou 0007, Fan Zhu 0001, Li Liu 0004, Ling Shao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Learning Multi-Attention Context Graph for Group-Based Re-IdentificationabstractLearning to re-identify or retrieve a group of people across non-overlapped camera systems has important applications in video surveillance. However, most existing methods focus on (single) person re-identification (re-id), ignoring the fact that people often walk in groups in real scenarios. In this work, we take a step further and consider employing context information for identifying groups of people, i.e., group re-id. On the one hand, group re-id is more challenging than single person re-id, since it requires both a robust modeling of local individual person appearance (with different illumination conditions, pose/viewpoint variations, and occlusions), as well as full awareness of global group structures (with group layout and group member variations). On the other hand, we believe that person re-id can be greatly enhanced by incorporating additional visual context from neighboring group members, a task which we formulate as group-aware (single) person re-id. In this paper, we propose a novel unified framework based on graph neural networks to simultaneously address the above two group-based re-id tasks, i.e., group re-id and group-aware person re-id. Specifically, we construct a context graph with group members as its nodes to exploit dependencies among different people. A multi-level attention mechanism is developed to formulate both intra-group and inter-group context, with an additional self-attention module for robust graph-level representations by attentively aggregating node-level features. The proposed model can be directly generalized to tackle group-aware person re-id using node-level representations. Meanwhile, to facilitate the deployment of deep learning models on these tasks, we build a new group re-id dataset which contains more than 3.8K images with 1.5K annotated groups, an order of magnitude larger than existing group re-id datasets. Extensive experiments on the novel dataset as well as three existing datasets clearly demonstrate the effectiveness of the proposed framework for both group-based re-id tasks. Yichao Yan, Jie Qin 0004, Bingbing Ni, Jiaxin Chen 0002, Li Liu 0004, Fan Zhu 0001, Wei-Shi Zheng 0001, Xiaokang Yang 0001, Ling Shao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2022 | A Generalized Method for Binary Optimization: Convergence Analysis and ApplicationsabstractBinary optimization problems (BOPs) arise naturally in many fields, such as information retrieval, computer vision, and machine learning. Most existing binary optimization methods either use continuous relaxation which can cause large quantization errors, or incorporate a highly specific algorithm that can only be used for particular loss functions. To overcome these difficulties, we propose a novel generalized optimization method, named Alternating Binary Matrix Optimization (ABMO), for solving BOPs. ABMO can handle BOPs with/without orthogonality or linear constraints for a large class of loss functions. ABMO involves rewriting the binary, orthogonality and linear constraints for BOPs as an intersection of two closed sets, then iteratively dividing the original problems into several small optimization problems that can be solved as closed forms. To provide a strict theoretical convergence analysis, we add a sufficiently small perturbation and translate the original problem to an approximated problem whose feasible set is continuous. We not only provide rigorous mathematical proof for the convergence to a stationary and feasible point, but also derive the convergence rate of the proposed algorithm. The promising results obtained from four binary optimization tasks validate the superiority and the generality of ABMO compared with the state-of-the-art methods. Huan Xiong, Mengyang Yu, Li Liu 0004, Fan Zhu 0001, Jie Qin 0004, Fumin Shen, Ling Shao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | Generalized Zero-Shot Learning With Multiple Graph Adaptive Generative NetworksabstractGenerative adversarial networks (GANs) for (generalized) zero-shot learning (ZSL) aim to generate unseen image features when conditioned on unseen class embeddings, each of which corresponds to one unique category. Most existing works on GANs for ZSL generate features by merely feeding the seen image feature/class embedding (combined with random Gaussian noise) pairs into the generator/discriminator for a two-player minimax game. However, the structure consistency of the distributions among the real/fake image features, which may shift the generated features away from their real distribution to some extent, is seldom considered. In this paper, to align the weights of the generator for better structure consistency between real/fake features, we propose a novel multigraph adaptive GAN (MGA-GAN). Specifically, a Wasserstein GAN equipped with a classification loss is trained to generate discriminative features with structure consistency. MGA-GAN leverages the multigraph similarity structures between sliced seen real/fake feature samples to assist in updating the generator weights in the local feature manifold. Moreover, correlation graphs for the whole real/fake features are adopted to guarantee structure correlation in the global feature manifold. Extensive evaluations on four benchmarks demonstrate well the superiority of MGA-GAN over its state-of-the-art counterparts. Guosen Xie, Zheng Zhang 0006, Guoshuai Liu, Fan Zhu 0001, Li Liu 0004, Ling Shao 0001, Xuelong Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2021 | Group Whitening: Balancing Learning Efficiency and Representational CapacityabstractBatch normalization (BN) is an important technique commonly incorporated into deep learning models to perform standardization within mini-batches. The merits of BN in improving a model’s learning efficiency can be further amplified by applying whitening, while its drawbacks in estimating population statistics for inference can be avoided through group normalization (GN). This paper proposes group whitening (GW), which exploits the advantages of the whitening operation and avoids the disadvantages of normalization within mini-batches. In addition, we analyze the constraints imposed on features by normalization, and show how the batch size (group number) affects the performance of batch (group) normalized networks, from the perspective of model’s representational capacity. This analysis provides theoretical guidance for applying GW in practice. Finally, we apply the proposed GW to ResNet and ResNeXt architectures and conduct experiments on the ImageNet and COCO benchmarks. Results show that GW consistently improves the performance of different architectures, with absolute gains of 1.02% ∼ 1.49% in top-1 accuracy on ImageNet and 1.82% ∼ 3.21% in bounding box AP on COCO. Lei Huang 0015, Yi Zhou 0007, Li Liu 0004, Fan Zhu 0001, Ling Shao 0001 |
CVPR | 4 |
| 2021 | Anchor-Free Person SearchabstractPerson search aims to simultaneously localize and identify a query person from realistic, uncropped images, which can be regarded as the unified task of pedestrian detection and person re-identification (re-id). Most existing works employ two-stage detectors like Faster-RCNN, yielding encouraging accuracy but with high computational overhead. In this work, we present the Feature-Aligned Person Search Network (AlignPS), the first anchor-free framework to efficiently tackle this challenging task. AlignPS explicitly addresses the major challenges, which we summarize as the misalignment issues in different levels (i.e., scale, region, and task), when accommodating an anchor-free detector for this task. More specifically, we propose an aligned feature aggregation module to generate more discriminative and robust feature embeddings by following a "re-id first" principle. Such a simple design directly improves the baseline anchor-free model on CUHK-SYSU by more than 20% in mAP. Moreover, AlignPS outperforms state-of-the-art two-stage methods, with a higher speed. The code is available at https://github.com/daodaofr/AlignPS. Yichao Yan, Jinpeng Li 0004, Jie Qin 0004, Song Bai 0001, Shengcai Liao, Li Liu 0004, Fan Zhu 0001, Ling Shao 0001 |
CVPR | 7 |
| 2021 | P2-Net: Joint Description and Detection of Local Features for Pixel and Point MatchingabstractAccurately describing and detecting 2D and 3D key-points is crucial to establishing correspondences across images and point clouds. Despite a plethora of learning-based 2D or 3D local feature descriptors and detectors having been proposed, the derivation of a shared descriptor and joint keypoint detector that directly matches pixels and points remains under-explored by the community. This work takes the initiative to establish fine-grained correspondences between 2D images and 3D point clouds. In order to directly match pixels and points, a dual fully-convolutional framework is presented that maps 2D and 3D inputs into a shared latent representation space to simultaneously describe and detect keypoints. Furthermore, an ultra-wide reception mechanism and a novel loss function are designed to mitigate the intrinsic information variations between pixel and point local regions. Extensive experimental results demonstrate that our framework shows competitive performance in fine-grained matching between images and point clouds and achieves state-of-the-art results for the task of indoor visual localization. Our source code is available at https://github.com/BingCS/P2-Net. Bing Wang 0013, Changhao Chen, Zhaopeng Cui, Jie Qin 0004, Xiaoxuan Lu 0001, Zhengdi Yu, Peijun Zhao, Zhen Dong 0005, Fan Zhu 0001, Agathoniki Trigoni, Andrew Markham |
ICCV | 9 |
| 2021 | Internal and external memory set containment join
Chengcheng Yang, Dong Deng 0001, Shuo Shang, Fan Zhu 0001, Li Liu 0004, Ling Shao 0001 |
VLDB J. | 4 |
| 2021 | Correction to: Internal and external memory set containment join
Chengcheng Yang, Dong Deng 0001, Shuo Shang, Fan Zhu 0001, Li Liu 0004, Ling Shao 0001 |
VLDB J. | 4 |
| 2020 | Controllable Orthogonalization in Training DNNsabstractOrthogonality is widely used for training deep neural networks (DNNs) due to its ability to maintain all singular values of the Jacobian close to 1 and reduce redundancy in representation. This paper proposes a computationally efficient and numerically stable orthogonalization method using Newton's iteration (ONI), to learn a layer-wise orthogonal weight matrix in DNNs. ONI works by iteratively stretching the singular values of a weight matrix towards 1. This property enables it to control the orthogonality of a weight matrix by its number of iterations. We show that our method improves the performance of image classification networks by effectively controlling the orthogonality to provide an optimal tradeoff between optimization benefits and representational capacity reduction. We also show that ONI stabilizes the training of generative adversarial networks (GANs) by maintaining the Lipschitz continuity of a network, similar to spectral normalization (SN), and further outperforms SN by providing controllable orthogonality. Lei Huang 0015, Li Liu 0004, Fan Zhu 0001, Diwen Wan, Zehuan Yuan, Bo Li 0026, Ling Shao 0001 |
CVPR | 3 |
| 2020 | An Investigation Into the Stochasticity of Batch WhiteningabstractBatch Normalization (BN) is extensively employed in various network architectures by performing standardization within mini-batches. A full understanding of the process has been a central target in the deep learning communities. Unlike existing works, which usually only analyze the standardization operation, this paper investigates the more general Batch Whitening (BW). Our work originates from the observation that while various whitening transformations equivalently improve the conditioning, they show significantly different behaviors in discriminative scenarios and training Generative Adversarial Networks (GANs). We attribute this phenomenon to the stochasticity that BW introduces. We quantitatively investigate the stochasticity of different whitening transformations and show that it correlates well with the optimization behaviors during training. We also investigate how stochasticity relates to the estimation of population statistics during inference. Based on our analysis, we provide a framework for designing and comparing BW algorithms in different scenarios. Our proposed BW algorithm improves the residual networks by a significant margin on ImageNet classification. Besides, we show that the stochasticity of BW can improve the GAN's performance with, however, the sacrifice of the training stability. Lei Huang 0015, Yi Zhou 0007, Fan Zhu 0001, Li Liu 0004, Ling Shao 0001 |
CVPR | 4 |
| 2020 | Auto-Encoding Twin-Bottleneck HashingabstractConventional unsupervised hashing methods usually take advantage of similarity graphs, which are either pre-computed in the high-dimensional space or obtained from random anchor points. On the one hand, existing methods uncouple the procedures of hash function learning and graph construction. On the other hand, graphs empirically built upon original data could introduce biased prior knowledge of data relevance, leading to sub-optimal retrieval performance. In this paper, we tackle the above problems by proposing an efficient and adaptive code-driven graph, which is updated by decoding in the context of an auto-encoder. Specifically, we introduce into our framework twin bottlenecks (i.e., latent variables) that exchange crucial information collaboratively. One bottleneck (i.e., binary codes) conveys the high-level intrinsic data structure captured by the code-driven graph to the other (i.e., continuous variables for low-level detail information), which in turn propagates the updated network feedback for the encoder to learn more discriminative binary codes. The auto-encoding learning objective literally rewards the code-driven graph to learn an optimal encoder. Moreover, the proposed model can be simply optimized by gradient descent without violating the binary constraints. Experiments on benchmarked datasets clearly show the superiority of our framework over the state-of-the-art hashing methods. Our source code can be found at https://github.com/ymcidence/TBH. Yuming Shen, Jie Qin 0004, Jiaxin Chen 0002, Mengyang Yu, Li Liu 0004, Fan Zhu 0001, Fumin Shen, Ling Shao 0001 |
CVPR | 6 |
| 2020 | Learning Multi-Granular Hypergraphs for Video-Based Person Re-IdentificationabstractVideo-based person re-identification (re-ID) is an important research topic in computer vision. The key to tackling the challenging task is to exploit both spatial and temporal clues in video sequences. In this work, we propose a novel graph-based framework, namely Multi-Granular Hypergraph (MGH), to pursue better representational capabilities by modeling spatiotemporal dependencies in terms of multiple granularities. Specifically, hypergraphs with different spatial granularities are constructed using various levels of part-based features across the video sequence. In each hypergraph, different temporal granularities are captured by hyperedges that connect a set of graph nodes (i.e., part-based features) across different temporal ranges. Two critical issues (misalignment and occlusion) are explicitly addressed by the proposed hypergraph propagation and feature aggregation schemes. Finally, we further enhance the overall video representation by learning more diversified graph-level representations of multiple granularities based on mutual information minimization. Extensive experiments on three widely-adopted benchmarks clearly demonstrate the effectiveness of the proposed framework. Notably, 90.0% top-1 accuracy on MARS is achieved using MGH, outperforming the state-of-the-arts. Yichao Yan, Jie Qin 0004, Jiaxin Chen 0002, Li Liu 0004, Fan Zhu 0001, Ying Tai, Ling Shao 0001 |
CVPR | 5 |
| 2020 | Learning Attentive and Hierarchical Representations for 3D Shape Recognition
Jiaxin Chen 0002, Jie Qin 0004, Yuming Shen, Li Liu 0004, Fan Zhu 0001, Ling Shao 0001 |
ECCV (15) | 5 |
| 2020 | Layer-Wise Conditioning Analysis in Exploring the Learning Dynamics of DNNs
Lei Huang 0015, Jie Qin 0004, Li Liu 0004, Fan Zhu 0001, Ling Shao 0001 |
ECCV (2) | 4 |
| 2020 | Invertible Zero-Shot Recognition Flows
Yuming Shen, Jie Qin 0004, Lei Huang 0015, Li Liu 0004, Fan Zhu 0001, Ling Shao 0001 |
ECCV (16) | 5 |
| 2020 | Region Graph Embedding Network for Zero-Shot Learning
Guosen Xie, Li Liu 0004, Fan Zhu 0001, Fang Zhao 0006, Zheng Zhang 0006, Yazhou Yao, Jie Qin 0004, Ling Shao 0001 |
ECCV (4) | 3 |
| 2020 | On the Number of Linear Regions of Convolutional Neural NetworksabstractOne fundamental problem in deep learning is understanding the outstanding performance of deep Neural Networks (NNs) in practice. One explanation for the superiority of NNs is that they can realize a large class of complicated functions, i.e., they have powerful expressivity. The expressivity of a ReLU NN can be quantified by the maximal number of linear regions it can separate its input space into. In this paper, we provide several mathematical results needed for studying the linear regions of CNNs, and use them to derive the maximal and average numbers of linear regions for one-layer ReLU CNNs. Furthermore, we obtain upper and lower bounds for the number of linear regions of multi-layer ReLU CNNs. Our results suggest that deeper CNNs have more powerful expressivity than their shallow counterparts, while CNNs have more expressivity than fully-connected NNs per parameter. Huan Xiong, Lei Huang 0015, Mengyang Yu, Li Liu 0004, Fan Zhu 0001, Ling Shao 0001 |
ICML | 5 |
| 2020 | Improved Residual Networks for Image and Video RecognitionabstractResidual networks (ResNets) represent a powerful type of convolutional neural network (CNN) architecture, widely adopted and used in various tasks. In this work we propose an improved version of ResNets. Our proposed improvements address all three main components of a ResNet: the flow of information through the network layers, the residual building block, and the projection shortcut. We are able to show consistent improvements in accuracy and learning convergence over the baseline. For instance, on ImageNet dataset, using the ResNet with 50 layers, for top-1 accuracy we can report a 1.19% improvement over the baseline in one setting and around 2% boost in another. Importantly, these improvements are obtained without increasing the model complexity. Our proposed approach allows us to train extremely deep networks, while the baseline shows severe optimization issues. We report results on three tasks over six datasets: image classification (ImageNet, CIFAR-10 and CIFAR-100), object detection (COCO) and video action recognition (Kinetics-400 and Something-Something-v2). In the deep learning era, we establish a new milestone for the depth of a CNN. We successfully train a 404-layer deep CNN on the ImageNet dataset and a 3002-layer network on CIFAR-10 and CIFAR-100, while the baseline is not able to converge at such extreme depths. Code and models are publicly available at: https://github.com/iduta/iresnet. I. C. Duta, Li Liu 0004, Fan Zhu 0001, Ling Shao 0001 |
ICPR | 3 |
| 2020 | Deep Local Binary Coding for Person Re-Identification by Delving into the DetailsabstractPerson re-identification (ReID) has recently received extensive research interests due to its diverse applications in multimedia analysis and computer vision. However, the majority of existing works focus on improving matching accuracy, while ignoring matching efficiency. In this work, we present a novel binary representation learning framework for efficient person ReID, namely Deep Local Binary Coding (DLBC). Different from existing deep binary ReID approaches, DLBC attempts to learn discriminative binary codes by explicitly interacting with local visual details. Specifically, DLBC first extracts a set of local features from spatially salient regions of pedestrian images. Subsequently, DLBC formulates a new binary-local semantic mutual information (BSMI) maximization term, based on which a self-lifting (SL) block is built to further exploit the semantic importance of local features. The BSMI term together with the SL block simultaneously enhances the dependency of binary codes on selected local features as well as their robustness to cross-view visual inconsistency. In addition, an efficient optimizing method is developed to train the proposed deep models with orthogonal and binary constraints. Extensive experiments reveal that DLBC significantly minimizes the accuracy gap between binary ReID methods and the state-of-the-art real-valued ones, whilst remarkably reducing query time and memory cost. Jiaxin Chen 0002, Jie Qin 0004, Yichao Yan, Lei Huang 0015, Li Liu 0004, Fan Zhu 0001, Ling Shao 0001 |
ACM Multimedia | 6 |
| 2020 | Projection based weight normalization: Efficient method for optimization on oblique manifold in DNNs
Lei Huang 0015, Xianglong Liu 0001, Jie Qin 0004, Fan Zhu 0001, Li Liu 0004, Ling Shao 0001 |
Pattern Recognit. | 4 |
| 2020 | Deep quantization generative networks
Diwen Wan, Fumin Shen, Li Liu 0004, Fan Zhu 0001, Lei Huang 0015, Mengyang Yu, Heng Tao Shen, Ling Shao 0001 |
Pattern Recognit. | 4 |
| 2020 | Zero-VAE-GAN: Generating Unseen Features for Generalized and Transductive Zero-Shot LearningabstractZero-shot learning (ZSL) is a challenging task due to the lack of unseen class data during training. Existing works attempt to establish a mapping between the visual and class spaces through a common intermediate semantic space. The main limitation of existing methods is the strong bias towards seen class, known as the domain shift problem, which leads to unsatisfactory performance in both conventional and generalized ZSL tasks. To tackle this challenge, we propose to convert ZSL to the conventional supervised learning by generating features for unseen classes. To this end, a joint generative model that couples variational autoencoder (VAE) and generative adversarial network (GAN), called Zero-VAE-GAN, is proposed to generate high-quality unseen features. To enhance the class-level discriminability, an adversarial categorization network is incorporated into the joint framework. Besides, we propose two self-training strategies to augment unlabeled unseen features for the transductive extension of our model, addressing the domain shift problem to a large extent. Experimental results on five standard benchmarks and a large-scale dataset demonstrate the superiority of our generative model over the state-of-the-art methods for conventional, especially generalized ZSL tasks. Moreover, the further improvement of the transductive setting demonstrates the effectiveness of the proposed self-training strategies. Xingsong Hou, Jie Qin 0004, Jiaxin Chen 0002, Li Liu 0004, Fan Zhu 0001, Zhao Zhang 0001, Ling Shao 0001 |
IEEE Trans. Image Process. | 6 |
| 2020 | NLH: A Blind Pixel-Level Non-Local Method for Real-World Image DenoisingabstractNon-local self similarity (NSS) is a powerful prior of natural images for image denoising. Most of existing denoising methods employ similar patches, which is a patch-level NSS prior. In this paper, we take one step forward by introducing a pixel-level NSS prior, i.e., searching similar pixels across a non-local region. This is motivated by the fact that finding closely similar pixels is more feasible than similar patches in natural images, which can be used to enhance image denoising performance. With the introduced pixel-level NSS prior, we propose an accurate noise level estimation method, and then develop a blind image denoising method based on the lifting Haar transform and Wiener filtering techniques. Experiments on benchmark datasets demonstrate that, the proposed method achieves much better performance than previous non-deep methods, and is still competitive with existing state-of-the-art deep learning based methods on real-world image denoising. The code is publicly available athttps://github.com/njusthyk1972/NLH. Yingkun Hou, Jun Xu 0019, Mingxia Liu 0001, Guanghai Liu 0001, Li Liu 0004, Fan Zhu 0001, Ling Shao 0001 |
IEEE Trans. Image Process. | 6 |
| 2020 | Noisy-as-Clean: Learning Self-Supervised Denoising From Corrupted ImageabstractSupervised deep networks have achieved promising performance on image denoising, by learning image priors and noise statistics on plenty pairs of noisy and clean images. Unsupervised denoising networks are trained with only noisy images. However, for an unseen corrupted image, both supervised and unsupervised networks ignore either its particular image prior, the noise statistics, or both. That is, the networks learned from external images inherently suffer from a domain gap problem: the image priors and noise statistics are very different between the training and test images. This problem becomes more clear when dealing with the signal dependent realistic noise. To circumvent this problem, in this work, we propose a novel "Noisy-As-Clean" (NAC) strategy of training self-supervised denoising networks. Specifically, the corrupted test image is directly taken as the "clean" target, while the inputs are synthetic images consisted of this corrupted image and a second yet similar corruption. A simple but useful observation on our NAC is: as long as the noise is weak, it is feasible to learn a self-supervised network only with the corrupted image, approximating the optimal parameters of a supervised network learned with pairs of noisy and clean images. Experiments on synthetic and realistic noise removal demonstrate that, the DnCNN and ResNet networks trained with our self-supervised NAC strategy achieve comparable or better performance than the original ones and previous supervised/unsupervised/self-supervised networks. The code is publicly available at https://github.com/csjunxu/Noisy-As-Clean. Jun Xu 0019, Ming-Ming Cheng, Li Liu 0004, Fan Zhu 0001, Zhou Xu 0003, Ling Shao 0001 |
IEEE Trans. Image Process. | 5 |
| 2020 | STAR: A Structure and Texture Aware Retinex ModelabstractRetinex theory is developed mainly to decompose an image into the illumination and reflectance components by analyzing local image derivatives. In this theory, larger derivatives are attributed to the changes in reflectance, while smaller derivatives are emerged in the smooth illumination. In this paper, we utilize exponentiated local derivatives (with an exponent γ) of an observed image to generate its structure map and texture map. The structure map is produced by been amplified with γ > 1, while the texture map is generated by been shrank with γ < 1. To this end, we design exponential filters for the local derivatives, and present their capability on extracting accurate structure and texture maps, influenced by the choices of exponents γ. The extracted structure and texture maps are employed to regularize the illumination and reflectance components in Retinex decomposition. A novel Structure and Texture Aware Retinex (STAR) model is further proposed for illumination and reflectance decomposition of a single image. We solve the STAR model by an alternating optimization algorithm. Each sub-problem is transformed into a vectorized least squares regression, with closed-form solutions. Comprehensive experiments on commonly tested datasets demonstrate that, the proposed STAR model produce better quantitative and qualitative performance than previous competing methods, on illumination and reflectance decomposition, low-light image enhancement, and color correction. The code is publicly available at https://github.com/csjunxu/STAR. Jun Xu 0019, Yingkun Hou, Dongwei Ren, Li Liu 0004, Fan Zhu 0001, Mengyang Yu, Haoqian Wang, Ling Shao 0001 |
IEEE Trans. Image Process. | 5 |
| 2020 | Towards Automatic Construction of Diverse, High-Quality Image DatasetsabstractThe availability of labeled image datasets has been shown critical for high-level image understanding, which continuously drives the progress of feature designing and models developing. However, constructing labeled image datasets is laborious and monotonous. To eliminate manual annotation, in this work, we propose a novel image dataset construction framework by employing multiple textual queries. We aim at collecting diverse and accurate images for given queries from the Web. Specifically, we formulate noisy textual queries removing and noisy images filtering as a multi-view and multi-instance learning problem separately. Our proposed approach not only improves the accuracy but also enhances the diversity of the selected images. To verify the effectiveness of our proposed approach, we construct an image dataset with 100 categories. The experiments show significant performance gains by using the generated data of our approach on several tasks, such as image classification, cross-dataset generalization, and object detection. The proposed method also consistently outperforms existing weakly supervised and web-supervised approaches. Yazhou Yao, Jian Zhang 0002, Fumin Shen, Li Liu 0004, Fan Zhu 0001, Dongxiang Zhang, Heng Tao Shen |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2020 | Approximate Kernel Selection via Matrix ApproximationabstractKernel selection is of fundamental importance for the generalization of kernel methods. This article proposes an approximate approach for kernel selection by exploiting the approximability of kernel selection and the computational virtue of kernel matrix approximation. We define approximate consistency to measure the approximability of the kernel selection problem. Based on the analysis of approximate consistency, we solve the theoretical problem of whether, under what conditions, and at what speed, the approximate criterion is close to the accurate one, establishing the foundations of approximate kernel selection. We introduce two selection criteria based on error estimation and prove the approximate consistency of the multilevel circulant matrix (MCM) approximation and Nyström approximation under these criteria. Under the theoretical guarantees of the approximate consistency, we design approximate algorithms for kernel selection, which exploits the computational advantages of the MCM and Nyström approximations to conduct kernel selection in a linear or quasi-linear complexity. We experimentally validate the theoretical results for the approximate consistency and evaluate the effectiveness of the proposed kernel selection algorithms. Lizhong Ding 0001, Shizhong Liao, Yong Liu 0018, Li Liu 0004, Fan Zhu 0001, Yazhou Yao, Ling Shao 0001, Xin Gao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2020 | SRSC: Selective, Robust, and Supervised Constrained Feature Representation for Image ClassificationabstractFeature representation learning, an emerging topic in recent years, has achieved great progress. Powerful learned features can lead to excellent classification accuracy. In this article, a selective and robust feature representation framework with a supervised constraint (SRSC) is presented. SRSC seeks a selective, robust, and discriminative subspace by transforming the original feature space into the category space. Particularly, we add a selective constraint to the transformation matrix (or classifier parameter) that can select discriminative dimensions of the input samples. Moreover, a supervised regularization is tailored to further enhance the discriminability of the subspace. To relax the hard zero-one label matrix in the category space, an additional error term is also incorporated into the framework, which can lead to a more robust transformation matrix. SRSC is formulated as a constrained least square learning (feature transforming) problem. For the SRSC problem, an inexact augmented Lagrange multiplier method (ALM) is utilized to solve it. Extensive experiments on several benchmark data sets adequately demonstrate the effectiveness and superiority of the proposed method. The proposed SRSC approach has achieved better performances than the compared counterpart methods. Guosen Xie, Zheng Zhang 0006, Li Liu 0004, Fan Zhu 0001, Xu-Yao Zhang, Ling Shao 0001, Xuelong Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2020 | Exploiting Web Images for Multi-Output Classification: From Category to SubcategoriesabstractStudies present that dividing categories into subcategories contributes to better image classification. Existing image subcategorization works relying on expert knowledge and labeled images are both time-consuming and labor-intensive. In this article, we propose to select and subsequently classify images into categories and subcategories. Specifically, we first obtain a list of candidate subcategory labels from untagged corpora. Then, we purify these subcategory labels through calculating the relevance to the target category. To suppress the search error and noisy subcategory label-induced outlier images, we formulate outlier images removing and the optimal classification models learning as a unified problem to jointly learn multiple classifiers, where the classifier for a category is obtained by combining multiple subcategory classifiers. Compared with the existing subcategorization works, our approach eliminates the dependence on expert knowledge and labeled images. Extensive experiments on image categorization and subcategorization demonstrate the superiority of our proposed approach. Yazhou Yao, Fumin Shen, Guosen Xie, Li Liu 0004, Fan Zhu 0001, Jian Zhang 0002, Heng Tao Shen |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2019 | Deep Sketch-Shape Hashing With Segmented 3D Stochastic ViewingabstractSketch-based 3D shape retrieval has been extensively studied in recent works, most of which focus on improving the retrieval accuracy, whilst neglecting the efficiency. In this paper, we propose a novel framework for efficient sketch-based 3D shape retrieval, i.e., Deep Sketch-Shape Hashing (DSSH), which tackles the challenging problem from two perspectives. Firstly, we propose an intuitive 3D shape representation method to deal with unaligned shapes with arbitrary poses. Specifically, the proposed Segmented Stochastic-viewing Shape Network models discriminative 3D representations by a set of 2D images rendered from multiple views, which are stochastically selected from non-overlapping spatial segments of a 3D sphere. Secondly, Batch-Hard Binary Coding (BHBC) is developed to learn semantics-preserving compact binary codes by mining the hardest samples. The overall framework is jointly learned by developing an alternating iteration algorithm. Extensive experimental results on three benchmarks show that DSSH improves both the retrieval efficiency and accuracy remarkably, compared to the state-of-the-art methods. Jiaxin Chen 0002, Jie Qin 0004, Li Liu 0004, Fan Zhu 0001, Fumin Shen, Jin Xie 0001, Ling Shao 0001 |
CVPR | 4 |
| 2019 | Iterative Normalization: Beyond Standardization Towards Efficient WhiteningabstractBatch Normalization (BN) is ubiquitously employed for accelerating neural network training and improving the generalization capability by performing standardization within mini-batches. Decorrelated Batch Normalization (DBN) further boosts the above effectiveness by whitening. However, DBN relies heavily on either a large batch size, or eigen-decomposition that suffers from poor efficiency on GPUs. We propose Iterative Normalization (IterNorm), which employs Newton’s iterations for much more efficient whitening, while simultaneously avoiding the eigen-decomposition. Furthermore, we develop a comprehensive study to show IterNorm has better trade-off between optimization and generalization, with theoretical and experimental support. To this end, we exclusively introduce Stochastic Normalization Disturbance (SND), which measures the inherent stochastic uncertainty of samples when applied to normalization operations. With the support of SND, we provide natural explanations to several phenomena from the perspective of optimization, e.g., why group-wise whitening of DBN generally outperforms full-whitening and why the accuracy of BN degenerates with reduced batch sizes. We demonstrate the consistently improved performance of IterNorm with extensive experiments on CIFAR-10 and ImageNet over BN and DBN. Lei Huang 0015, Yi Zhou 0007, Fan Zhu 0001, Li Liu 0004, Ling Shao 0001 |
CVPR | 3 |
| 2019 | Building Detail-Sensitive Semantic Segmentation Networks With Polynomial PoolingabstractSemantic segmentation is an important computer vision task, which aims to allocate a semantic label to each pixel in an image. When training a segmentation model, it is common to fine-tune a classification network pre-trained on a large-scale dataset. However, as an intrinsic property of the classification model, invariance to spatial perturbation resulting from the lose of detail-sensitivity prevents segmentation networks from achieving high performance. The use of standard poolings is one of the key factors for this invariance. The most common standard poolings are max and average pooling. Max pooling can increase both the invariance to spatial perturbations and the non-linearity of the networks. Average pooling, on the other hand, is sensitive to spatial perturbations, but is a linear function. For semantic segmentation, we prefer both the preservation of detailed cues within a local feature region and non-linearity that increases a network's functional complexity. In this work, we propose a polynomial pooling (P-pooling) function that finds an intermediate form between max and average pooling to provide an optimally balanced and self-adjusted pooling strategy for semantic segmentation. The P-pooling is differentiable and can be applied into a variety of pre-trained networks. Extensive studies on the PASCAL VOC, Cityscapes and ADE20k datasets demonstrate the superiority of P-pooling over other poolings. Experiments on various network architectures and state-of-the-art training strategies also show that models with P-pooling layers consistently outperform those directly fine-tuned using pre-trained classification models. Zhen Wei 0001, Jingyi Zhang 0005, Li Liu 0004, Fan Zhu 0001, Fumin Shen, Yi Zhou 0007, Si Liu 0001, Yao Sun 0004, Ling Shao 0001 |
CVPR | 4 |
| 2019 | Attentive Region Embedding Network for Zero-Shot LearningabstractZero-shot learning (ZSL) aims to classify images from unseen categories, by merely utilizing seen class images as the training data. Existing works on ZSL mainly leverage the global features or learn the global regions, from which, to construct the embeddings to the semantic space. However, few of them study the discrimination power implied in local image regions (parts), which, in some sense, correspond to semantic attributes, have stronger discrimination than attributes, and can thus assist the semantic transfer between seen/unseen classes. In this paper, to discover (semantic) regions, we propose the attentive region embedding network (AREN), which is tailored to advance the ZSL task. Specifically, AREN is end-to-end trainable and consists of two network branches, i.e., the attentive region embedding (ARE) stream, and the attentive compressed second-order embedding (ACSE) stream. ARE is capable of discovering multiple part regions under the guidance of the attention and the compatibility loss. Moreover, a novel adaptive thresholding mechanism is proposed for suppressing redundant (such as background) attention regions. To further guarantee more stable semantic transfer from the perspective of second-order collaboration, ACSE is incorporated into the AREN. In the comprehensive evaluations on four benchmarks, our models achieve state-of-the-art performances under ZSL setting, and compelling results under generalized ZSL setting. Guosen Xie, Li Liu 0004, Xiao-Bo Jin, Fan Zhu 0001, Zheng Zhang 0006, Jie Qin 0004, Yazhou Yao, Ling Shao 0001 |
CVPR | 4 |
| 2019 | Collaborative Learning of Semi-Supervised Segmentation and Classification for Medical ImagesabstractMedical image analysis has two important research areas: disease grading and fine-grained lesion segmentation. Although the former problem often relies on the latter, the two are usually studied separately. Disease severity grading can be treated as a classification problem, which only requires image-level annotations, while the lesion segmentation requires stronger pixel-level annotations. However, pixel-wise data annotation for medical images is highly time-consuming and requires domain experts. In this paper, we propose a collaborative learning method to jointly improve the performance of disease grading and lesion segmentation by semi-supervised learning with an attention mechanism. Given a small set of pixel-level annotated data, a multi-lesion mask generation model first performs the traditional semantic segmentation task. Then, based on initially predicted lesion maps for large quantities of image-level annotated data, a lesion attentive disease grading model is designed to improve the severity classification accuracy. Meanwhile, the lesion attention model can refine the lesion maps using class-specific information to fine-tune the segmentation model in a semi-supervised manner. An adversarial architecture is also integrated for training. With extensive experiments on a representative medical problem called diabetic retinopathy (DR), we validate the effectiveness of our method and achieve consistent improvements over state-of-the-art methods on three public datasets. Yi Zhou 0007, Xiaodong He 0004, Lei Huang 0015, Li Liu 0004, Fan Zhu 0001, Shanshan Cui, Ling Shao 0001 |
CVPR | 5 |
| 2019 | RANet: Ranking Attention Network for Fast Video Object SegmentationabstractDespite online learning (OL) techniques have boosted the performance of semi-supervised video object segmentation (VOS) methods, the huge time costs of OL greatly restricts their practicality. Matching based and propagation based methods run at a faster speed by avoiding OL techniques. However, they are limited by sub-optimal accuracy, due to mismatching and drifting problems. In this paper, we develop a real-time yet very accurate Ranking Attention Network (RANet) for VOS. Specifically, to integrate the insights of matching based and propagation based methods, we employ an encoder-decoder framework to learn pixel-level similarity and segmentation in an end-to-end manner. To better utilize the similarity maps, we propose a novel ranking attention module, which automatically ranks and selects these maps for fine-grained VOS performance. Experiments on DAVIS16 and DAVIS17 datasets show that our RANet achieves the best speed-accuracy trade-off, e.g., with 33 milliseconds per frame and J&F=85.5% on DAVIS16. With OL, our RANet reaches J&F=87.1% on DAVIS16, exceeding state-of-the-art VOS methods. The code can be found at https://github.com/Storife/RANet. Ziqin Wang, Jun Xu 0019, Li Liu 0004, Fan Zhu 0001, Ling Shao 0001 |
ICCV | 4 |
| 2019 | Relational Attention Network for Crowd CountingabstractCrowd counting is receiving rapidly growing research interests due to its potential application value in numerous real-world scenarios. However, due to various challenges such as occlusion, insufficient resolution and dynamic backgrounds, crowd counting remains an unsolved problem in computer vision. Density estimation is a popular strategy for crowd counting, where conventional density estimation methods perform pixel-wise regression without explicitly accounting the interdependence of pixels. As a result, independent pixel-wise predictions can be noisy and inconsistent. In order to address such an issue, we propose a Relational Attention Network (RANet) with a self-attention mechanism for capturing interdependence of pixels. The RANet enhances the self-attention mechanism by accounting both short-range and long-range interdependence of pixels, where we respectively denote these implementations as local self-attention (LSA) and global self-attention (GSA). We further introduce a relation module to fuse LSA and GSA to achieve more informative aggregated feature representations. We conduct extensive experiments on four public datasets, including ShanghaiTech A, ShanghaiTech B, UCF-CC-50 and UCF-QNRF. Experimental results on all datasets suggest RANet consistently reduces estimation errors and surpasses the state-of-the-art approaches by large margins. Zehao Xiao, Fan Zhu 0001, Xiantong Zhen, Xianbin Cao 0001, Ling Shao 0001 |
ICCV | 4 |
| 2019 | Attentional Neural Fields for Crowd CountingabstractCrowd counting has recently generated huge popularity in computer vision, and is extremely challenging due to the huge scale variations of objects. In this paper, we propose the Attentional Neural Field (ANF) for crowd counting via density estimation. Within the encoder-decoder network, we introduce conditional random fields (CRFs) to aggregate multi-scale features, which can build more informative representations. To better model pair-wise potentials in CRFs, we incorperate non-local attention mechanism implemented as inter- and intra-layer attentions to expand the receptive field to the entire image respectively within the same layer and across different layers, which captures long-range dependencies to conquer huge scale variations. The CRFs coupled with the attention mechanism are seamlessly integrated into the encoder-decoder network, establishing an ANF that can be optimized end-to-end by back propagation. We conduct extensive experiments on four public datasets, including ShanghaiTech, WorldEXPO 10, UCF-CC-50 and UCF-QNRF. The results show that our ANF achieves high counting performance, surpassing most previous methods. Fan Zhu 0001, Xiantong Zhen, Xianbin Cao 0001, Ling Shao 0001 |
ICCV | 4 |
| 2019 | LCJoin: Set Containment Join via List CrosscuttingabstractA set containment join operates on two set-valued attributes with a subset (⊆) relationship as the join condition. It has many real-world applications, such as in publish/subscribe services and inclusion dependency discovery. Existing solutions can be broadly classified into union-oriented and intersection-oriented methods. Based on several recent studies, union-oriented methods are not competitive as they involve an expensive subset enumeration step. Intersection-oriented methods build an inverted index on one attribute and perform inverted list intersection on another attribute. Existing intersection-oriented methods intersect inverted lists one-by-one. In contrast, in this paper, we propose to intersect all the inverted lists simultaneously while skipping many irrelevant entries in the lists. To share computation, we utilize the prefix tree structure and extend our novel list intersection method to operate on the prefix tree. To further improve the efficiency, we propose to partition the data and use different methods to process each partition. We evaluated our methods using both real-world and synthetic datasets. Experimental results show that our approach outperforms existing methods by up to 10×. Dong Deng 0001, Chengcheng Yang, Shuo Shang, Fan Zhu 0001, Li Liu 0004, Ling Shao 0001 |
ICDE | 4 |
| 2019 | Toward Efficient Navigation of Massive-Scale Geo-Textual StreamsabstractWith the popularization of portable devices, numerous applications continuously produce huge streams of geo-tagged textual data, thus posing challenges to index geo-textual streaming data efficiently, which is an important task in both data management and AI applications, e.g., real-time data streams mining and targeted advertising. This, however, is not possible with the state-of-the-art indexing methods as they focus on search optimizations of static datasets, and have high index maintenance cost. In this paper, we present NQ-tree, which combines new structure designs and self-tuning methods to navigate between update and search efficiency. Our contributions include: (1) the design of multiple stores each with a different emphasis on write-friendness and read-friendness; (2) utilizing data compression techniques to reduce the I/O cost; (3) exploiting both spatial and keyword information to improve the pruning efficiency; (4) proposing an analytical cost model, and using an online self-tuning method to achieve efficient accesses to different workloads. Experiments on two real-world datasets show that NQ-tree outperforms two well designed baselines by up to 10×. Chengcheng Yang, Lisi Chen 0001, Shuo Shang, Fan Zhu 0001, Li Liu 0004, Ling Shao 0001 |
IJCAI | 4 |
| 2019 | Dynamically Visual Disambiguation of Keyword-based Image SearchabstractDue to the high cost of manual annotation, learning directly from the web has attracted broad attention. One issue that limits their performance is the problem of visual polysemy. To address this issue, we present an adaptive multi-model framework that resolves polysemy by visual disambiguation. Compared to existing methods, the primary advantage of our approach lies in that our approach can adapt to the dynamic changes in the search results. Our proposed framework consists of two major steps: we first discover and dynamically select the text queries according to the image search results, then we employ the proposed saliency-guided deep multi-instance learning network to remove outliers and learn classification models for visual disambiguation. Extensive experiments demonstrate the superiority of our proposed approach. Yazhou Yao, Zeren Sun, Fumin Shen, Li Liu 0004, Limin Wang 0002, Fan Zhu 0001, Lizhong Ding 0001, Gangshan Wu, Ling Shao 0001 |
IJCAI | 6 |
| 2019 | High-Resolution Diabetic Retinopathy Image Synthesis Manipulated by Grading and Lesions
Yi Zhou 0007, Xiaodong He 0004, Shanshan Cui, Fan Zhu 0001, Li Liu 0004, Ling Shao 0001 |
MICCAI (1) | 4 |
| 2019 | Generative Reconstructive Hashing for Incomplete Video AnalysisabstractIn the literature of video analysis, most researches, such as retrieval and recognition, hypothesize that each input video contains at least one complete semantic entity, e.g. an activity, action and event.However, this hypothesis does not hold in many realistic scenarios due to two main reasons. First, complete videos whose qualities are good enough for automatic analysis are not always accessible because of heavy motion blur, occlusions, interruptions, etc. % Second, extracting features from complete videos always fails to meet up with speed and storage requirements in large-scale use cases.To tackle these challenges, incomplete videos are more useful, but researches on them are seldom mentioned. In this paper, we propose a novel and effective hashing framework specialized in large-scale incomplete video analysis called Generative Reconstructive Hashing (GRH). To begin with, an adversarial generative network that is specially designed to map incomplete video features to the feature distributions of complete videos, so that features of incomplete videos become indistinguishable from those of complete videos. Then, the discriminative hashing module further fills the gap between full video features and estimated features from partial videos by projecting both features into a common binary feature space, which allows improvement in efficiency compared with real-value based methods. GRH is the first end-to-end framework for incomplete video analysis. Extensive experiments on various datasets demonstrate GRH's superior effectiveness and efficiency on retrieval and recognition tasks. GRH outperforms the recent state-of-the-art methods by 5.44/3.22/4.82 in terms of MAPs on HMDB51/UCF101/CCV datasets, respectively. Jingyi Zhang 0005, Zhen Wei 0001, I. C. Duta, Fumin Shen, Li Liu 0004, Fan Zhu 0001, Xing Xu 0001, Ling Shao 0001, Heng Tao Shen |
ACM Multimedia | 6 |
| 2019 | Two Generator Game: Learning to Sample via Linear Goodness-of-Fit TestabstractLearning the probability distribution of high-dimensional data is a challenging problem. To solve this problem, we formulate a deep energy adversarial network (DEAN), which casts the energy model learned from real data into an optimization of a goodness-of-fit (GOF) test statistic. DEAN can be interpreted as a GOF game between two generative networks, where one explicit generative network learns an energy-based distribution that fits the real data, and the other implicit generative network is trained by minimizing a GOF test statistic between the energy-based distribution and the generated data, such that the underlying distribution of the generated data is close to the energy-based distribution. We design a two-level alternative optimization procedure to train the explicit and implicit generative networks, such that the hyper-parameters can also be automatically learned. Experimental results show that DEAN achieves high quality generations compared to the state-of-the-art approaches. Lizhong Ding 0001, Mengyang Yu, Li Liu 0004, Fan Zhu 0001, Yong Liu 0018, Yu Li 0006, Ling Shao 0001 |
NeurIPS | 4 |
| 2018 | Structure-Aware 3D Shape Synthesis from Single-View Images
Xuyang Hu, Fan Zhu 0001, Li Liu 0004, Jin Xie 0001, Jun Tang 0007, Nian Wang 0002, Fumin Shen, Ling Shao 0001 |
BMVC | 2 |
| 2018 | TBN: Convolutional Neural Network with Ternary Inputs and Binary Weights
Diwen Wan, Fumin Shen, Li Liu 0004, Fan Zhu 0001, Jie Qin 0004, Ling Shao 0001, Heng Tao Shen |
ECCV (2) | 4 |
| 2018 | Highly-Economized Multi-view Binary Compression for Scalable Image Clustering
Zheng Zhang 0006, Li Liu 0004, Jie Qin 0004, Fan Zhu 0001, Fumin Shen, Yong Xu 0001, Ling Shao 0001, Heng Tao Shen |
ECCV (12) | 4 |
| 2018 | Generative Domain-Migration Hashing for Sketch-to-Image Retrieval
Jingyi Zhang 0005, Fumin Shen, Li Liu 0004, Fan Zhu 0001, Mengyang Yu, Ling Shao 0001, Heng Tao Shen, Luc Van Gool |
ECCV (2) | 4 |
| 2018 | Learning to Synthesize 3D Indoor Scenes from Monocular ImagesabstractDepth images have always been playing critical roles for indoor scene understanding problems, and are particularly important for tasks in which 3D inferences are involved. However, since depth images are not universally available, abandoning them from the testing stage can significantly improve the generality of a method. In this work, we consider the scenarios where depth images are not available in the testing data, and propose to learn a convolutional long short-term memory (Conv LSTM) network and a regression convolutional neural network (regression ConvNet) using only monocular RGB images. The proposed networks benefit from 2D segmentations, object-level spatial context, object-scene dependencies and objects' geometric information, where optimization is governed by the semantic label loss, which measures the label consistencies of both objects and scenes, and the 3D geometrical loss, which measures the correctness of objects' 6Dof estimation. Conv LSTM and regression ConvNet are applied to scene/object classification, object detection and 6Dof estimation tasks respectively, where we utilize the joint inference from both networks and further provide the perspective of synthesizing fully rigged 3D scenes according to objects' arrangements in monocular images. Both quantitative and qualitative experimental results are provided on the NYU-v2 dataset, and we demonstrate that the proposed Conv LSTM can achieve state-of-the-art performance without requiring the depth information. Fan Zhu 0001, Li Liu 0004, Jin Xie 0001, Fumin Shen, Ling Shao 0001, Yi Fang 0006 |
ACM Multimedia | 1 |
| 2018 | Face recognition with a small occluded training set using spatial and statistical pooling
Yang Long 0001, Fan Zhu 0001, Ling Shao 0001, Junwei Han 0001 |
Inf. Sci. | 2 |
| 2018 | Deep Nonlinear Metric Learning for 3-D Shape RetrievalabstractEffective 3-D shape retrieval is an important problem in 3-D shape analysis. Recently, feature learning-based shape retrieval methods have been widely studied, where the distance metrics between 3-D shape descriptors are usually hand-crafted. In this paper, motivated by the fact that deep neural network has the good ability to model nonlinearity, we propose to learn an effective nonlinear distance metric between 3-D shape descriptors for retrieval. First, the locality-constrained linear coding method is employed to encode each vertex on the shape and the encoding coefficient histogram is formed as the global 3-D shape descriptor to represent the shape. Then, a novel deep metric network is proposed to learn a nonlinear transformation to map the 3-D shape descriptors to a nonlinear feature space. The proposed deep metric network minimizes a discriminative loss function that can enforce the similarity between a pair of samples from the same class to be small and the similarity between a pair of samples from different classes to be large. Finally, the distance between the outputs of the metric network is used as the similarity for shape retrieval. The proposed method is evaluated on the McGill, SHREC'10 ShapeGoogle, and SHREC'14 Human shape datasets. Experimental results on the three datasets validate the effectiveness of the proposed method. Jin Xie 0001, Guoxian Dai, Fan Zhu 0001, Ling Shao 0001, Yi Fang 0006 |
IEEE Trans. Cybern. | 3 |
| 2017 | Deep Correlated Metric Learning for Sketch-based 3D Shape RetrievalabstractThe explosive growth of 3D models has led to the pressing demand for an efficient searching system. Traditional model-based search is usually not convenient, since people don't always have 3D model available by side. The sketch-based 3D shape retrieval is a promising candidate due to its simpleness and efficiency. The main challenge for sketch-based 3D shape retrieval is the discrepancy across different domains. In the paper, we propose a novel deep correlated metric learning (DCML) method to mitigate the discrepancy between sketch and 3D shape domains. The proposed DCML trains two distinct deep neural networks (one for each domain) jointly with one loss, which learns two deep nonlinear transformations to map features from both domains into a nonlinear feature space. The proposed loss, including discriminative loss and correlation loss, aims to increase the discrimination of features within each domain as well as the correlation between different domains. In the transfered space, the discriminative loss minimizes the intra-class distance of the deep transformed features and maximizes the inter-class distance of the deep transformed features at least a predefined margin within each domain, while the correlation loss focuses on minimizing the distribution discrepancy across different domains. Our proposed method is evaluated on SHREC 2013 and 2014 benchmarks, and the experimental results demonstrate the superiority of our proposed method over the state-of-the-art methods. Guoxian Dai, Jin Xie 0001, Fan Zhu 0001, Yi Fang 0006 |
AAAI | 3 |
| 2017 | Learning Barycentric Representations of 3D Shapes for Sketch-Based 3D Shape RetrievalabstractRetrieving 3D shapes with sketches is a challenging problem since 2D sketches and 3D shapes are from two heterogeneous domains, which results in large discrepancy between them. In this paper, we propose to learn barycenters of 2D projections of 3D shapes for sketch-based 3D shape retrieval. Specifically, we first use two deep convolutional neural networks (CNNs) to extract deep features of sketches and 2D projections of 3D shapes. For 3D shapes, we then compute the Wasserstein barycenters of deep features of multiple projections to form a barycentric representation. Finally, by constructing a metric network, a discriminative loss is formulated on the Wasserstein barycenters of 3D shapes and sketches in the deep feature space to learn discriminative and compact 3D shape and sketch features for retrieval. The proposed method is evaluated on the SHREC13 and SHREC14 sketch track benchmark datasets. Compared to the state-of-the-art methods, our proposed method can significantly improve the retrieval performance. Jin Xie 0001, Guoxian Dai, Fan Zhu 0001, Yi Fang 0006 |
CVPR | 3 |
| 2017 | DeepShape: Deep-Learned Shape Descriptor for 3D Shape RetrievalabstractComplex geometric variations of 3D models usually pose great challenges in 3D shape matching and retrieval. In this paper, we propose a novel 3D shape feature learning method to extract high-level shape features that are insensitive to geometric deformations of shapes. Our method uses a discriminative deep auto-encoder to learn deformation-invariant shape features. First, a multiscale shape distribution is computed and used as input to the auto-encoder. We then impose the Fisher discrimination criterion on the neurons in the hidden layer to develop a deep discriminative auto-encoder. Finally, the outputs from the hidden layers of the discriminative auto-encoders at different scales are concatenated to form the shape descriptor. The proposed method is evaluated on four benchmark datasets that contain 3D models with large geometric variations: McGill, SHREC'10 ShapeGoogle, SHREC'14 Human and SHREC'14 Large Scale Comprehensive Retrieval Track Benchmark datasets. Experimental results on the benchmark datasets demonstrate the effectiveness of the proposed method for 3D shape retrieval. Jin Xie 0001, Guoxian Dai, Fan Zhu 0001, Edward K. Wong, Yi Fang 0006 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2017 | Progressive Shape-Distribution-Encoder for Learning 3D Shape RepresentationabstractSince there are complex geometric variations with 3D shapes, extracting efficient 3D shape features is one of the most challenging tasks in shape matching and retrieval. In this paper, we propose a deep shape descriptor by learning shape distributions at different diffusion time via a progressive shape-distribution-encoder (PSDE). First, we develop a shape distribution representation with the kernel density estimator to characterize the intrinsic geometry structures of 3D shapes. Then, we propose to learn a deep shape feature through an unsupervised PSDE. Specially, the unsupervised PSDE aims at modeling the complex non-linear transform of the estimated shape distributions between consecutive diffusion time. In order to characterize the intrinsic structures of 3D shapes more efficiently, we stack multiple PSDEs to form a network structure. Finally, we concatenate all neurons in the middle hidden layers of the unsupervised PSDE network to form an unsupervised shape descriptor for retrieval. Furthermore, by imposing an additional constraint on the outputs of all hidden layers, we propose a supervised PSDE to form a supervised shape descriptor. For each hidden layer, the similarity between a pair of outputs from the same class is as large as possible and the similarity between a pair of outputs from different classes is as small as possible. The proposed method is evaluated on three benchmark 3D shape data sets with large geometric variations, i.e., McGill, SHREC'10 ShapeGoogle, and SHREC'14 Human data sets, and the experimental results demonstrate the superiority of the proposed method to the existing approaches. Jin Xie 0001, Fan Zhu 0001, Guoxian Dai, Ling Shao 0001, Yi Fang 0006 |
IEEE Trans. Image Process. | 2 |
| 2016 | Learning Cross-Domain Neural Networks for Sketch-Based 3D Shape RetrievalabstractSketch-based 3D shape retrieval, which returns a set of relevant 3D shapes based on users' input sketch queries, has been receiving increasing attentions in both graphics community and vision community. In this work, we address the sketch-based 3D shape retrieval problem with a novel Cross-Domain Neural Networks (CDNN) approach, which is further extended to Pyramid Cross-Domain Neural Networks (PCDNN) by cooperating with a hierarchical structure. In order to alleviate the discrepancies between sketch features and 3D shape features, a neural network pair that forces identical representations at the target layer for instances of the same class is trained for sketches and 3D shapes respectively. By constructing cross-domain neural networks at multiple pyramid levels, a many-to-one relationship is established between a 3D shape feature and sketch features extracted from different scales. We evaluate the effectiveness of both CDNN and PCDNN approach on the extended large-scale SHREC 2014 benchmark and compare with some other well established methods. Experimental results suggest that both CDNN and PCDNN can outperform state-of-the-art performance, where PCDNN can further improve CDNN when employing a hierarchical structure. Fan Zhu 0001, Jin Xie 0001, Yi Fang 0006 |
AAAI | 1 |
| 2016 | Heat Diffusion Long-Short Term Memory Learning for 3D Shape Analysis
Fan Zhu 0001, Jin Xie 0001, Yi Fang 0006 |
ECCV (7) | 1 |
| 2016 | Color object recognition via cross-domain learning on RGB-D imagesabstractThis paper addresses the object recognition problem using multiple-domain inputs. We present a novel approach that utilizes labeled RGB-D data in the training stage, where depth features are extracted for enhancing the discriminative capability of the original learning system that only relies on RGB images. The highly dissimilar source and target domain data are mapped into a unified feature space through transfer at both feature and classifier levels. In order to alleviate cross-domain discrepancy, we employ a state-of-the-art domain-adaptive dictionary learning algorithm that updates image representations in both domains and the classifier parameters simultaneously. The proposed method is trained on a RGB-D Object dataset and evaluated on the Caltech-256 dataset. Experimental results suggest that our approach can lead to significant performance gain over the state-of-the-art methods. Yawen Huang, Fan Zhu 0001, Ling Shao 0001, Alejandro F. Frangi |
ICRA | 2 |
| 2016 | Recognising occluded multi-view actions using local nearest neighbour embedding
Yang Long 0001, Fan Zhu 0001, Ling Shao 0001 |
Comput. Vis. Image Underst. | 2 |
| 2016 | Dual many-to-one-encoder-based transfer learning for cross-dataset human action recognition
Fan Zhu 0001, Edward K. Wong, Yi Fang 0006 |
Image Vis. Comput. | 2 |
| 2016 | From handcrafted to learned representations for human action recognition: A survey
Fan Zhu 0001, Ling Shao 0001, Jin Xie 0001, Yi Fang 0006 |
Image Vis. Comput. | 1 |
| 2016 | Learning a discriminative deformation-invariant 3D shape descriptor via many-to-one encoder
Guoxian Dai, Jin Xie 0001, Fan Zhu 0001, Yi Fang 0006 |
Pattern Recognit. Lett. | 3 |
| 2016 | Linear discrimination dictionary learning for shape descriptors
Meng Wang 0001, Jin Xie 0001, Fan Zhu 0001, Yi Fang 0006 |
Pattern Recognit. Lett. | 3 |
| 2015 | 3D deep shape descriptorabstractShape descriptor is a concise yet informative representation that provides a 3D object with an identification as a member of some category. We have developed a concise deep shape descriptor to address challenging issues from ever-growing 3D datasets in areas as diverse as engineering, medicine, and biology. Specifically, in this paper, we developed novel techniques to extract concise but geometrically informative shape descriptor and new methods of defining Eigen-shape descriptor and Fisher-shape descriptor to guide the training of a deep neural network. Our deep shape descriptor tends to maximize the inter-class margin while minimize the intra-class variance. Our new shape descriptor addresses the challenges posed by the high complexity of 3D model and data representation, and the structural variations and noise present in 3D models. Experimental results on 3D shape retrieval demonstrate the superior performance of deep shape descriptor over other state-of-the-art techniques in handling noise, incompleteness and structural variations. Yi Fang 0006, Jin Xie 0001, Guoxian Dai, Meng Wang 0001, Fan Zhu 0001, Edward K. Wong |
CVPR | 5 |
| 2015 | Deepshape: Deep learned shape descriptor for 3D shape matching and retrievalabstractComplex geometric structural variations of 3D model usually pose great challenges in 3D shape matching and retrieval. In this paper, we propose a high-level shape feature learning scheme to extract features that are insensitive to deformations via a novel discriminative deep auto-encoder. First, a multiscale shape distribution is developed for use as input to the auto-encoder. Then, by imposing the Fisher discrimination criterion on the neurons in the hidden layer, we developed a novel discriminative deep auto-encoder for shape feature learning. Finally, the neurons in the hidden layers from multiple discriminative auto-encoders are concatenated to form a shape descriptor for 3D shape matching and retrieval. The proposed method is evaluated on the representative datasets that contain 3D models with large geometric variations, i.e., Mcgill and SHREC'10 ShapeGoogle datasets. Experimental results on the benchmark datasets demonstrate the effectiveness of the proposed method for 3D shape matching and retrieval. Jin Xie 0001, Yi Fang 0006, Fan Zhu 0001, Edward K. Wong |
CVPR | 3 |
| 2015 | Progressive Shape-Distribution-Encoder for 3D Shape RetrievalabstractIn this paper, we propose a deep shape descriptor by learning the shape distributions at different diffusion time via a progressive deep shape-distribution-encoder. First, we develop a shape distribution representation with the kernel density estimator to characterize the intrinsic geometrical structure of the shape. Then, we propose to learn discriminative shape features through a progressive shape-distribution-encoder. Specially, the progressive shape-distribution-encoder aims at modeling the complex non-linear transform of the estimated shape distributions between consecutive diffusion time. Furthermore, in order to characterize the intrinsic structure of the shape more efficiently, we stack multiple proposed progressive shape-distribution-encoders to form a neural network structure. Finally, we concatenated all neurons in the hidden layers of the progressive shape-distribution-encoder network to form a discriminative shape descriptor for retrieval. The proposed method is evaluated on three benchmark 3D shape datasets %with large geometric variations, i.e., McGill, SHREC'10 ShapeGoogle and SHREC'14 Human datasets, and the experimental results demonstrate the superiority of our method to the existing approaches. Jin Xie 0001, Fan Zhu 0001, Guoxian Dai, Yi Fang 0006 |
ACM Multimedia | 2 |
| 2015 | Learning Pairwise Neural Network Encoder for Depth Image-based 3D Model RetrievalabstractWith the emergence of RGB-D cameras (e.g., Kinect), the sensing capability of artificial intelligence systems has been dramatically increased, and as a consequence, a wide range of depth image-based human-machine interaction applications are proposed. In design industry, a 3D model always contains abundant information, which are required for manufacture. Since depth images can be conveniently acquired, a retrieval system that can return 3D models based on depth image inputs can assist or improve the traditional product design process. In this work, we address the depth image-based 3D model retrieval problem. By extending the neural network to a neural network pair with identical output layers for objects of the same category, unified domain-invariant representations can be learned based on the low-level mismatched depth image features and 3D model features. A unique advantage of the framework is that the correspondence information between depth images and 3D models are not required, so that it can easily be generalized to large-scale databases. In order to evaluate the effectiveness of our approach, depth images (with Kinect-type noise) in the NYU Depth V2 dataset are used as queries to retrieve 3D models of the same categories in the SHREC 2014 dataset. Experimental results suggest that our approach can outperform the state-of-the-arts methods, and the paradigm that directly uses the original representations of depth images and 3D models for retrieval. Jing Zhu 0002, Fan Zhu 0001, Edward K. Wong, Yi Fang 0006 |
ACM Multimedia | 2 |
| 2015 | Transfer Learning for Visual Categorization: A SurveyabstractRegular machine learning and data mining techniques study the training data for future inferences under a major assumption that the future data are within the same feature space or have the same distribution as the training data. However, due to the limited availability of human labeled training data, training data that stay in the same feature space or have the same distribution as the future data cannot be guaranteed to be sufficient enough to avoid the over-fitting problem. In real-world applications, apart from data in the target domain, related data in a different domain can also be included to expand the availability of our prior knowledge about the target future data. Transfer learning addresses such cross-domain learning problems by extracting useful information from data in a related domain and transferring them for being used in target tasks. In recent years, with transfer learning being applied to visual categorization, some typical problems, e.g., view divergence in action recognition tasks and concept drifting in image classification tasks, can be efficiently solved. In this paper, we survey state-of-the-art transfer learning algorithms in visual categorization applications, such as object recognition, image classification, and human action recognition. Ling Shao 0001, Fan Zhu 0001, Xuelong Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2014 | Boosted Cross-Domain Categorization
Fan Zhu 0001, Ling Shao 0001, Jun Tang 0007 |
BMVC | 1 |
| 2014 | Cross-Modality Submodular Dictionary Learning for Information RetrievalabstractThis paper addresses the problem of joint modeling of multimedia components in different media forms. We consider the information retrieval task across both text and image documents, which includes retrieving relevant images that closely match the description in a text query and retrieving text documents that best explain the content of an image query. A greedy dictionary construction approach is introduced for learning an isomorphic feature space, to which cross-modality data can be adapted while data smoothness is guaranteed. The proposed objective function consists of two reconstruction error terms for both modalities and a Maximum Mean Discrepancy (MMD) term that measures the cross-modality discrepancy. Optimization of the reconstruction terms and the MMD term yields a compact and modality-adaptive dictionary pair. We formulate the joint combinatorial optimization problem by maximizing variance reduction over a candidate signal set while constraining the dictionary size and coefficients' sparsity. By exploiting the submodularity and the monotonicity property of the proposed objective function, the optimization problem can be solved by a highly efficient greedy algorithm, and is guaranteed to be at least a (e - 1)=/e≈0.632- approximation to the optimum. The proposed method achieves state-of-the-art performance on the Wikipedia dataset. Fan Zhu 0001, Ling Shao 0001, Mengyang Yu |
CIKM | 1 |
| 2014 | Submodular Object RecognitionabstractWe present a novel object recognition framework based on multiple figure-ground hypotheses with a large object spatial support, generated by bottom-up processes and mid-level cues in an unsupervised manner. We exploit the benefit of regression for discriminating segments' categories and qualities, where a regressor is trained to each category using the overlapping observations between each figure-ground segment hypothesis and the ground-truth of the target category in an image. Object recognition is achieved by maximizing a submodular objective function, which maximizes the similarities between the selected segments (i.e., facility locations) and their group elements (i.e., clients), penalizes the number of selected segments, and more importantly, encourages the consistency of object categories corresponding to maximum regression values from different category-specific regressors for the selected segments. The proposed framework achieves impressive recognition results on three benchmark datasets, including PASCAL VOC 2007, Caltech-101 and ETHZ-shape. Fan Zhu 0001, Zhuolin Jiang, Ling Shao 0001 |
CVPR | 1 |
| 2014 | Correspondence-Free Dictionary Learning for Cross-View Action RecognitionabstractIn this paper, we propose a novel unsupervised approach to enabling sparse action representations for the cross-view action recognition scenario. Superior to previous cross-view action recognition methods, neither target view label information nor correspondence annotations are required in this approach. Low-level dense trajectory action features are first coded according to their feature space localities within the same view by projecting each descriptor into its local-coordinate system under the locality constraints. Actions across each pair of views are additionally decomposed to sparse linear combinations of basis atoms, a.k.a., dictionary elements, which are learned to reconstruct the original data while simultaneously forcing similar actions to have identical representations in an unsupervised manner. Consequently, cross-view knowledge is retained through the learned basis atoms, so that high-level representations of actions from both views can be considered to possess the same data distribution and can be directly fed into the classifier. The proposed approach achieves improved performance compared to state-of-the-art methods on the multi-view IXMAS data set, and leads to a new experimental setting that is closer to real-world applications. Fan Zhu 0001, Ling Shao 0001 |
ICPR | 1 |
| 2014 | Weakly-Supervised Cross-Domain Dictionary Learning for Visual Recognition
Fan Zhu 0001, Ling Shao 0001 |
Int. J. Comput. Vis. | 1 |
| 2013 | Enhancing Action Recognition by Cross-Domain Dictionary LearningabstractWe present a novel cross-dataset action recognition framework that utilizes relevant actions from other visual domains as auxiliary knowledge for enhancing the learning sys-tem in the target domain. The data distribution of relevant actions from a source dataset is adapted to match the data distribution of actions in the target dataset via a cross-domain discriminative dictionary learning method, through which a reconstructive, discrimina-tive and domain-adaptive dictionary-pair can be learned. Using selected categories from the HMDB51 dataset as the source domain actions, the proposed framework achieves outstanding performance on the UCF YouTube dataset. Fan Zhu 0001, Ling Shao 0001 |
BMVC | 1 |
| 2013 | Multi-view action recognition using local similarity random forests and sensor fusion
Fan Zhu 0001, Ling Shao 0001, Mingxiu Lin |
Pattern Recognit. Lett. | 1 |
| 2012 | One shot learning gesture recognition with Kinect sensorabstractGestures are both natural and intuitive for Human-Computer-Interaction (HCI) and the one-shot learning scenario is one of the real world situations in terms of gesture recognition problems. In this demo, we present a hand gesture recognition system using the Kinect sensor, which addresses the problem of one-shot learning gesture recognition with a user-defined training and testing system. Such a system can behave like a remote control where the user can allocate a specific function using a prefered gesture by performing it only once. To adopt the gesture recognition framework, the system first automatically segments an action sequence into atomic tokens, and then adopts the Extended-Motion-History-Image (Extended-MHI) for motion feature representation. We evaluate the performance of our system quantitatively in Chalearn Gesture Challenge, and apply it to a virtual one shot learning gesture recognition system. Di Wu 0009, Fan Zhu 0001, Ling Shao 0001, Hui Zhang 0062 |
ACM Multimedia | 2 |