Jie Lin 0001

dblp:88/6731-1 · DBLP profile ↗
← Back
75ranked-venue papers
7as first author
29since 2021 · last 2026
0000-0002-8971-0660ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 57 · 7 first-author · 15 since 2021Artificial intelligence and machine learning · 21 · 15 since 2021Databases, data management, data science and information retrieval · 9 · 2 first-author · 1 since 2021Systems, architecture and hardware · 7 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 LB-PTQ: Effective Low-Bit Post-Training Quantization for Vision Transformers
abstract
Recently, Vision Transformers (ViTs) have become the state-of-the-art architecture on various computer vision tasks including image classification, object detection and semantic segmentation. However, such success in high-accuracy performance comes at the price of high computational complexity, with typically tens of millions of or even more parameters in a Vision Transformer (ViT) model. Such a large volume of parameters makes it very difficult to deploy ViT models on mobile devices and cumbers their applications. In this paper, we present a novel post-training quantization approach that is able to quantize ViT models to very low bit widths, without the need of re-training. Prior works on post-training quantization for ViTs optimize the quantization of each layer separately thus leading to sub-optimal results. In contrast, we propose a unified learning framework that jointly optimizes the quantization of all layers to directly reduce the overall output error of the network. Moreover, we explore an important property of ViTs, i.e., the additivity property, revealing that the output error caused by the quantization of multiple layers equals the sum of the output error due to the quantization of each layer. Utilizing this property, we present a very efficient algorithm to solve the joint optimization problem with linear time complexity. We performed extensive experiments on the large-scale ImageNet dataset to evaluate the effectiveness of our approach. Empirical results show that our approach improves state-of-the-art noticeably on various ViT models and lowers the bit width from 8-bit to 6-bit without hurting the accuracy. Specifically, at 4 bits, our approach significantly outperforms existing works by 1.72%, 11.49%, 6.15%, and 3.54% on ViT-S, ViT-B, DeiT-S, and DeiT-B, respectively. In the end, we evaluate the performance when deploying our quantized models on hardware. Our approach achieves $1.5\times $ to $1.7\times $ speedups for the inference on NVIDIA A100 GPU.
Zhe Wang 0019, Kaixin Xu, Xue Geng, Jie Lin 0001, Mohamed M. Sabry, Min Wu 0008, Xiaoli Li 0001, Weisi Lin
IEEE Trans. Image Process.4
2025 Efficient Distortion-Minimized Layerwise Pruning
abstract
In this paper, we propose a post-training pruning framework that jointly optimizes layerwise pruning to minimize model output distortion. Through theoretical and empirical analysis, we discover an important additivity property of output distortion from pruning weights/channels in DNNs. Leveraging this property, we reformulate pruning optimization as a combinatorial problem and solve it with dynamic programming, achieving linear time complexity and making the algorithm very fast on CPUs. Furthermore, we optimize additivity-derived distortions using Hessian-based Taylor approximation to enhance pruning efficiency, accompanied by fine-grained complexity reduction techniques. Our method is evaluated on various DNN architectures, including CNNs, ViTs, and object detectors, and on vision tasks such as image classification on CIFAR-10 and ImageNet, and 3D object detection and various datasets. We achieve SoTA with significant FLOPs reductions without accuracy loss. Specifically, on CIFAR-10, we achieve up to $27.9\times$27.9×, $29.2\times$29.2×, and $14.9\times$14.9× FLOPs reductions on ResNet-32, VGG-16, and DenseNet-121, respectively. On ImageNet, we observe no accuracy loss with $1.69\times$1.69× and $2\times$2× FLOPs reductions on ResNet-50 and DeiT-Base, respectively. For 3D object detection, we achieve $\mathbf {3.89}\times, \mathbf {3.72}\times$3.89×,3.72× FLOPs reductions on CenterPoint and PVRCNN models. These results demonstrate the effectiveness and practicality of our approach for improving model performance through layer-adaptive weight pruning.
Kaixin Xu, Zhe Wang 0019, Runtao Huang, Xue Geng, Jie Lin 0001, Xulei Yang, Min Wu 0008, Xiaoli Li 0001, Weisi Lin
IEEE Trans. Pattern Anal. Mach. Intell.5
2025 PSRR-MaxpoolNMS++: Fast Non-Maximum Suppression With Discretization and Pooling
abstract
Non-maximum suppression (NMS) is an essential post-processing step for object detection. The de-facto standard for NMS, namely GreedyNMS, is not parallelizable and could thus be the performance bottleneck in object detection pipelines. MaxpoolNMS is introduced as a fast and parallelizable alternative to GreedyNMS. However, MaxpoolNMS is only capable of replacing the GreedyNMS at the first stage of two-stage detectors like Faster R-CNN. To address this issue, we observe that MaxpoolNMS employs the process of box coordinate discretization followed by local score argmax calculation, to discard the nested-loop pipeline in GreedyNMS to enable parallelizable implementations. In this paper, we introduce a simple Relationship Recovery module and a Pyramid Shifted MaxpoolNMS module to improve the above two stages, respectively. With these two modules, our PSRR-MaxpoolNMS is a generic and parallelizable approach, which can completely replace GreedyNMS at all stages in all detectors. Furthermore, we extend PSRR-MaxpoolNMS to the more powerful PSRR-MaxpoolNMS++. As for box coordinate discretization, we propose Density-based Discretization for better adherence to the target density of the suppression. As for local score argmax calculation, we propose an Adjacent Scale Pooling scheme for mining out the duplicated box pairs more accurately and efficiently. Extensive experiments demonstrate that both our PSRR-MaxpoolNMS and PSRR-MaxpoolNMS++ outperform MaxpoolNMS by a large margin. Additionally, PSRR-MaxpoolNMS++ not only surpasses PSRR-MaxpoolNMS but also attains competitive accuracy and much better efficiency when compared with GreedyNMS. Therefore, PSRR-MaxpoolNMS++ is a parallelizable NMS solution that can effectively replace GreedyNMS at all stages in all detectors.
Tianyi Zhang 0004, Chunyun Chen, Yun Liu 0011, Xue Geng, Mohamed M. Sabry, Jie Lin 0001
IEEE Trans. Pattern Anal. Mach. Intell.6
2025 From Algorithm to Hardware: A Survey on Efficient and Safe Deployment of Deep Neural Networks
abstract
Deep neural networks (DNNs) have been widely used in many artificial intelligence (AI) tasks. However, deploying them brings significant challenges due to the huge cost of memory, energy, and computation. To address these challenges, researchers have developed various model compression techniques such as model quantization and model pruning. Recently, there has been a surge in research on compression methods to achieve model efficiency while retaining performance. Furthermore, more and more works focus on customizing the DNN hardware accelerators to better leverage the model compression techniques. In addition to efficiency, preserving security and privacy is critical for deploying DNNs. However, the vast and diverse body of related works can be overwhelming. This inspires us to conduct a comprehensive survey on recent research toward the goal of high-performance, cost-efficient, and safe deployment of DNNs. Our survey first covers the mainstream model compression techniques, such as model quantization, model pruning, knowledge distillation, and optimizations of nonlinear operations. We then introduce recent advances in designing hardware accelerators that can adapt to efficient model compression approaches. In addition, we discuss how homomorphic encryption can be integrated to secure DNN deployment. Finally, we discuss several issues, such as hardware evaluation, generalization, and integration of various compression approaches. Overall, we aim to provide a big picture of efficient DNNs from algorithm to hardware accelerators and security perspectives.
Xue Geng, Zhe Wang 0019, Chunyun Chen, Qing Xu 0015, Kaixin Xu, Jin Chao, Manas Gupta, Xulei Yang, Zhenghua Chen, Mohamed M. Sabry, Jie Lin 0001, Min Wu 0008, Xiaoli Li 0001
IEEE Trans. Neural Networks Learn. Syst.11
2024 LPViT: Low-Power Semi-structured Pruning for Vision Transformers
Kaixin Xu, Zhe Wang 0019, Chunyun Chen, Xue Geng, Jie Lin 0001, Xulei Yang, Min Wu 0008, Xiaoli Li 0001, Weisi Lin
ECCV (71)5
2024 Evolving filter criteria for randomly initialized network pruning in image classification
Chenjing Liu, Peng Hu 0002, Jie Lin 0001, Yunhong Gong, Yingke Chen, Dezhong Peng, Xue Geng
Neurocomputing4
2024 LCReg: Long-tailed image classification with Latent Categories based Recognition
Weide Liu, Henghui Ding, Fayao Liu, Jie Lin 0001, Guosheng Lin
Pattern Recognit.6
2024 Deep Supervised Multi-View Learning With Graph Priors
abstract
This paper presents a novel method for supervised multi-view representation learning, which projects multiple views into a latent common space while preserving the discrimination and intrinsic structure of each view. Specifically, an apriori discriminant similarity graph is first constructed based on labels and pairwise relationships of multi-view inputs. Then, view-specific networks progressively map inputs to common representations whose affinity approximates the constructed graph. To achieve graph consistency, discrimination, and cross-view invariance, the similarity graph is enforced to meet the following constraints: 1) pairwise relationship should be consistent between the input space and common space for each view; 2) within-class similarity is larger than any between-class similarity for each view; 3) the inter-view samples from the same (or different) classes are mutually similar (or dissimilar). Consequently, the intrinsic structure and discrimination are preserved in the latent common space using an apriori approximation schema. Moreover, we present a sampling strategy to approach a sub-graph sampled from the whole similarity structure instead of approximating the graph of the whole dataset explicitly, thus benefiting lower space complexity and the capability of handling large-scale multi-view datasets. Extensive experiments show the promising performance of our method on five datasets by comparing it with 18 state-of-the-art methods.
Peng Hu 0002, Liangli Zhen, Xi Peng 0001, Hongyuan Zhu 0002, Jie Lin 0001, Xu Wang 0028, Dezhong Peng
IEEE Trans. Image Process.5
2024 On Representation Knowledge Distillation for Graph Neural Networks
abstract
Knowledge distillation (KD) is a learning paradigm for boosting resource-efficient graph neural networks (GNNs) using more expressive yet cumbersome teacher models. Past work on distillation for GNNs proposed the local structure preserving (LSP) loss, which matches local structural relationships defined over edges across the student and teacher's node embeddings. This article studies whether preserving the global topology of how the teacher embeds graph data can be a more effective distillation objective for GNNs, as real-world graphs often contain latent interactions and noisy edges. We propose graph contrastive representation distillation (G-CRD), which uses contrastive learning to implicitly preserve global topology by aligning the student node embeddings to those of the teacher in a shared representation space. Additionally, we introduce an expanded set of benchmarks on large-scale real-world datasets where the performance gap between teacher and student GNNs is non-negligible. Experiments across four datasets and 14 heterogeneous GNN architectures show that G-CRD consistently boosts the performance and robustness of lightweight GNNs, outperforming LSP (and a global structure preserving (GSP) variant of LSP) as well as baselines from 2-D computer vision. An analysis of the representational similarity among teacher and student embedding spaces reveals that G-CRD balances preserving local and global relationships, while structure preserving approaches are best at preserving one or the other.
Chaitanya K. Joshi, Fayao Liu, Xu Xun, Jie Lin 0001, Chuan-Sheng Foo
IEEE Trans. Neural Networks Learn. Syst.4
2023 Unsupervised Contrastive Cross-Modal Hashing
abstract
In this paper, we study how to make unsupervised cross-modal hashing (CMH) benefit from contrastive learning (CL) by overcoming two challenges. To be exact, i) to address the performance degradation issue caused by binary optimization for hashing, we propose a novel momentum optimizer that performs hashing operation learnable in CL, thus making on-the-shelf deep cross-modal hashing possible. In other words, our method does not involve binary-continuous relaxation like most existing methods, thus enjoying better retrieval performance; ii) to alleviate the influence brought by false-negative pairs (FNPs), we propose a Cross-modal Ranking Learning loss (CRL) which utilizes the discrimination from all instead of only the hard negative pairs, where FNP refers to the within-class pairs that were wrongly treated as negative pairs. Thanks to such a global strategy, CRL endows our method with better performance because CRL will not overuse the FNPs while ignoring the true-negative pairs. To the best of our knowledge, the proposed method could be one of the first successful contrastive hashing methods. To demonstrate the effectiveness of the proposed method, we carry out experiments on five widely-used datasets compared with 13 state-of-the-art methods. The code is available at https://github.com/penghu-cs/UCCH.
Peng Hu 0002, Hongyuan Zhu 0002, Jie Lin 0001, Dezhong Peng, Yin-Ping Zhao, Xi Peng 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2022 Point Discriminative Learning for Data-efficient 3D Point Cloud Analysis
abstract
3D point cloud analysis has drawn a lot of research attention due to its wide applications. However, collecting massive labelled 3D point cloud data is both time-consuming and labor-intensive. This calls for data-efficient learning methods. In this work we propose PointDisc, a point discriminative learning method to leverage self-supervisions for data-efficient 3D point cloud classification and segmentation. PointDisc imposes a novel point discrimination loss on the middle and global level features produced by the backbone network. This point discrimination loss enforces learned features to be consistent with points belonging to the corresponding local shape region and inconsistent with randomly sampled noisy points. We conduct extensive experiments on 3D object classification, 3D semantic and part segmentation, showing the benefits of PointDisc for data-efficient learning. Detailed analysis demonstrate that PointDisc learns unsupervised features that well capture local and global geometry.
Fayao Liu, Guosheng Lin, Chuan-Sheng Foo, Chaitanya K. Joshi, Jie Lin 0001
3DV5
2022 Scalable Hardware Acceleration of Non-Maximum Suppression
abstract
Non-maximum Suppression (NMS) in one- and two-stage object detection deep neural networks (e.g., SSD and Faster-RCNN) is becoming the computation bottleneck. In this paper, we introduce a hardware acceleration for the scalable PSRR-MaxpoolNMS algorithm. Our architecture shows 75.0× and 305× speedups compared to the software implementation of the PSRR-MaxpoolNMS as well as the hardware implementations of GreedyNMS, respectively, while simultaneously achieving comparable Mean Average Precision (mAP) to software-based floating-point implementations. Our architecture is 13.4× faster than the state-of-the-art NMS one. Our accelerator supports both one- and two-stage detectors, while supporting very high input resolutions (i.e., FHD)—essential input size for better detection accuracy.
Chunyun Chen, Tianyi Zhang 0004, Zehui Yu, Adithi Raghuraman, Shwetalaxmi Udayan, Jie Lin 0001, Mohamed M. Sabry
DATE6
2022 RDO-Q: Extremely Fine-Grained Channel-Wise Quantization via Rate-Distortion Optimization
Zhe Wang 0019, Jie Lin 0001, Xue Geng, Mohamed M. Sabry, Vijay Chandrasekhar 0001
ECCV (12)2
2022 Channel-Wise Bit Allocation for Deep Visual Feature Quantization
abstract
Intermediate deep visual feature compression and transmission is an emerging research topic, which enables a good balance among computing load, bandwidth usage and generalization ability for AI-based visual analysis in edge-cloud collaboration. Quantization and the corresponding rate-distortion optimization are the key techniques in deep feature compression. In this paper, by exploring the feature statistics and a greedy iterative algorithm, we propose a channel-wise bit allocation method for deep feature quantization optimizing for network output error. Given the limited rate and computational power, the proposed method can quantize features with small information loss. Moreover, the method also provides the option to handle the trade-offs between computational cost and quantization performance. Experimental results on ResNet and VGGNet features demonstrate the effectiveness of the proposed bit allocation method.
Wei Wang 0283, Zhuo Chen 0006, Zhe Wang 0019, Jie Lin 0001, Long Xu 0001, Weisi Lin
ICIP4
2022 Towards high performance homomorphic encryption for inference tasks on CPU: An MPI approach
Souhail Meftah, Benjamin Hong Meng Tan, Khin Mi Mi Aung, Yuxiao Lu, Jie Lin 0001, Bharadwaj Veeravalli
Future Gener. Comput. Syst.5
2022 QLP: Deep Q-Learning for Pruning Deep Neural Networks
abstract
We present a novel, deep Q-learning based method, QLP, for pruning deep neural networks (DNNs). Given a DNN, our method intelligently determines favorable layer-wise sparsity ratios, which are then implemented via unstructured, magnitude-based, weight pruning. In contrast to previous reinforcement learning (RL) based pruning methods, our method is not forced to prune a DNN within a single, sequential pass from the first layer to the last. It visits each layer multiple times and prunes them little by little at each visit, achieving superior granular pruning. Moreover, our method is not restricted to a subset of actions within the feasible action space. It has the flexibility to execute a whole range of sparsity ratios (0% - 100%) for each layer. This enables aggressive pruning without compromising accuracy. Furthermore, our method does not require a complex state definition; it features a simple, generic definition that is composed of only the index and the density of the layers, which leads to less computational demand while observing the state at each interaction. Lastly, our method utilizes a carefully designed curriculum that enables learning targeted policies for each sparsity regime, which helps to deliver better accuracy, especially at high sparsity levels. We conduct batched performance tests at compelling sparsity levels (up to 98%), present extensive ablation studies to justify our RL-related design choices, and compare our method with the state-of-the-art, including RL-based and other pruning methods. Our method sets the new state-of-the-art results in most of the experiments with ResNet-32 and ResNet-56 over CIFAR-10 dataset as well as ResNet-50 and MobileNet-v1 over ILSVRC2012 (ImageNet) dataset.
Efe Camci, Manas Gupta, Min Wu 0008, Jie Lin 0001
IEEE Trans. Circuits Syst. Video Technol.4
2022 Deep Semisupervised Multiview Learning With Increasing Views
abstract
In this article, we study two challenging problems in semisupervised cross-view learning. On the one hand, most existing methods assume that the samples in all views have a pairwise relationship, that is, it is necessary to capture or establish the correspondence of different views at the sample level. Such an assumption is easily isolated even in the semisupervised setting wherein only a few samples have labels that could be used to establish the correspondence. On the other hand, almost all existing multiview methods, including semisupervised ones, usually train a model using a fixed dataset, which cannot handle the data of increasing views. In practice, the view number will increase when new sensors are deployed. To address the above two challenges, we propose a novel method that employs multiple independent semisupervised view-specific networks (ISVNs) to learn representation for multiple views in a view-decoupling fashion. The advantages of our method are two-fold. Thanks to our specifically designed autoencoder and pseudolabel learning paradigm, our method shows an effective way to utilize both the labeled and unlabeled data while relaxing the data assumption of the pairwise relationship, that is, correspondence. Furthermore, with our view decoupling strategy, the proposed ISVNs could be separately trained, thus efficiently handling the data of increasing views without retraining the entire model. To the best of our knowledge, our ISVN could be one of the first attempts to make handling increasing views in the semisupervised setting possible, as well as an effective solution to the noncorresponding problem. To verify the effectiveness and efficiency of our method, we conduct comprehensive experiments by comparing 13 state-of-the-art approaches on four multiview datasets in terms of retrieval and classification.
Peng Hu 0002, Xi Peng 0001, Hongyuan Zhu 0002, Liangli Zhen, Jie Lin 0001, Huaibai Yan, Dezhong Peng
IEEE Trans. Cybern.5
2022 $A^3$-FKG: Attentive Attribute-Aware Fashion Knowledge Graph for Outfit Preference Prediction
abstract
With the booming development of the online fashion industry, effective personalized recommender systems have become indispensable for the convenience they brought to the customers and the profits to the e-commercial platforms. Estimating the user’s preference towards the outfit is at the core of a personalized recommendation system. Existing works on fashion recommendation are largely centering on modelling the clothing compatibility without considering the user factor or characterizing the user’s preference over the single item. However, how to effectively model the outfits with either few or even none interactions, is yet under-explored. In this paper, we address the task of personalized outfit preference prediction via a novelAttentiveAttribute-AwareFashionKnowledgeGraph ($A^3$-FKG), which is incorporated to build the association between different outfits with both outfit- and item- level attributes. Additionally, a two-level attention mechanism is developed to capture the user’s preference: 1) User-specific relation-aware attention layer, which captures the user’s fine-grained preferences with different focus on relations for learning outfit representation; 2) Target-aware attention layer, which characterizes the user’s latent diverse interests from his/her behavior sequences for learning user representation. Extensive experiments conducted on a large-scale fashion outfit dataset demonstrate significant improvements over other methods, which verify the excellence of our proposed framework.
Huijing Zhan, Jie Lin 0001, Kenan E. Ak, Boxin Shi, Ling-Yu Duan, Alex Chichung Kot
IEEE Trans. Multim.2
2021 OPQ: Compressing Deep Neural Networks with One-shot Pruning-Quantization
abstract
As Deep Neural Networks (DNNs) usually are overparameterized and have millions of weight parameters, it is challenging to deploy these large DNN models on resource-constrained hardware platforms, e.g., smartphones. Numerous network compression methods such as pruning and quantization are proposed to reduce the model size significantly, of which the key is to find suitable compression allocation (e.g., pruning sparsity and quantization codebook) of each layer. Existing solutions obtain the compression allocation in an iterative/manual fashion while finetuning the compressed model, thus suffering from the efficiency issue. Different from the prior art, we propose a novel One-shot Pruning-Quantization (OPQ) in this paper, which analytically solves the compression allocation with pre-trained weight parameters only. During finetuning, the compression module is fixed and only weight parameters are updated. To our knowledge, OPQ is the first work that reveals pre-trained model is sufficient for solving pruning and quantization simultaneously, without any complex iterative/manual optimization at the finetuning stage. Furthermore, we propose a unified channel-wise quantization method that enforces all channels of each layer to share a common codebook, which leads to low bit-rate allocation without introducing extra overhead brought by traditional channel-wise quantization. Comprehensive experiments on ImageNet with AlexNet/MobileNet-V1/ResNet-50 show that our method improves accuracy and training efficiency while obtains significantly higher compression rates compared to the state-of-the-art.
Peng Hu 0002, Xi Peng 0001, Hongyuan Zhu 0002, Mohamed M. Sabry, Jie Lin 0001
AAAI5
2021 Learning Cross-Modal Retrieval With Noisy Labels
abstract
Recently, cross-modal retrieval is emerging with the help of deep multimodal learning. However, even for unimodal data, collecting large-scale well-annotated data is expensive and time-consuming, and not to mention the additional challenges from multiple modalities. Although crowd-sourcing annotation, e.g., Amazon’s Mechanical Turk, can be utilized to mitigate the labeling cost, but leading to the unavoidable noise in labels for the non-expert annotating. To tackle the challenge, this paper presents a general Multi-modal Robust Learning framework (MRL) for learning with multimodal noisy labels to mitigate noisy samples and correlate distinct modalities simultaneously. To be specific, we propose a Robust Clustering loss (RC) to make the deep networks focus on clean samples instead of noisy ones. Besides, a simple yet effective multimodal loss function, called Multimodal Contrastive loss (MC), is proposed to maxi-mize the mutual information between different modalities, thus alleviating the interference of noisy samples and cross-modal discrepancy. Extensive experiments are conducted on four widely-used multimodal datasets to demonstrate the effectiveness of the proposed approach by comparing to 14 state-of-the-art methods.
Peng Hu 0002, Xi Peng 0001, Hongyuan Zhu 0002, Liangli Zhen, Jie Lin 0001
CVPR5
2021 PSRR-MaxpoolNMS: Pyramid Shifted MaxpoolNMS With Relationship Recovery
abstract
Non-maximum Suppression (NMS) is an essential post-processing step in modern convolutional neural networks for object detection. Unlike convolutions which are inherently parallel, the de-facto standard for NMS, namely GreedyNMS, cannot be easily parallelized and thus could be the performance bottleneck in convolutional object detection pipelines. MaxpoolNMS is introduced as a parallelizable alternative to GreedyNMS, which in turn enables faster speed than GreedyNMS at comparable accuracy. However, MaxpoolNMS is only capable of replacing the GreedyNMS at the first stage of two-stage detectors like Faster-RCNN. There is a significant drop in accuracy when applying MaxpoolNMS at the final detection stage, due to the fact that MaxpoolNMS fails to approximate GreedyNMS precisely in terms of bounding box selection. In this paper, we propose a general, parallelizable and configurable approach PSRR-MaxpoolNMS, to completely replace GreedyNMS at all stages in all detectors. By introducing a simple Relationship Recovery module and a Pyramid Shifted MaxpoolNMS module, our PSRR-MaxpoolNMS is able to approximate GreedyNMS more precisely than MaxpoolNMS. Comprehensive experiments show that our approach outperforms MaxpoolNMS by a large margin, and it is proven faster than GreedyNMS with comparable accuracy. For the first time, PSRR-MaxpoolNMS provides a fully parallelizable solution for customized hardware design, which can be reused for accelerating NMS everywhere.
Tianyi Zhang 0004, Jie Lin 0001, Peng Hu 0002, Mohamed M. Sabry
CVPR2
2021 Efficient Tunstall Decoder for Deep Neural Network Compression
abstract
Power and area-efficient deep neural network (DNN) designs are key in edge applications. Compact DNNs, via compression or quantization, enable such designs by significantly reducing memory footprint. Lossless entropy coding can further reduce the size of networks. It is then critical to provide hardware support for such entropy coding module to fully benefit from the resulted reduced memory requirement. In this work, we introduce Tunstall coding to compress the quantized weights. Tunstall coding can achieve high compression ratio as well as very fast decoding speed on various deep networks. We present two hardware-accelerated decoding techniques that provide streamlined decoding capabilities. We synthesize these designs targeting on FPGA. Results show that we achieve up to 6× faster decoding time versus state-of-the-art decoding methods.
Chunyun Chen, Zhe Wang 0019, Jie Lin 0001, Mohamed M. Sabry
DAC4
2021 Rate-Distortion Optimized Coding for Efficient CNN Compression
abstract
In this paper, we present a coding framework for deep convolutional neural network compression. Our approach utilizes the classical coding theories and formulates the compression of deep convolutional neural networks as a rate-distortion optimization problem. We incorporate three coding ingredients in the coding framework, including bit allocation, dead zone quantization, and Tunstall coding, to improve the rate-distortion frontier without noticeable system-level overhead introduced. Experimental results show that our approach achieves state-of-the-art results on various deep convolutional neural networks and obtains considerable speedup on two deep learning accelerators. Specifically, our approach achieves 20× compression ratio on ResNet-18, ResNet-34, and ResNet-50, and 10× compression ratio on the compact already model MobileNet-v2, without hurting the accuracy. We then examine the system level impact of our approach when deploying the compressed models to hardware platforms. Hardware simulation results show that our approach obtains up to 4.3× and 2.8× inference speedup on state-of-the-art deep learning accelerators TPU and Eyeriss, respectively.
Wang Zhe, Jie Lin 0001, Mohamed M. Sabry, Sean I. Young, Vijay Chandrasekhar 0001, Bernd Girod
DCC2
2021 PAN: Personalized Attention Network For Outfit Recommendation
abstract
Recent years have witnessed the dramatic development of e-fashion industry, it becomes essential to build an intelligent fashion recommender system. Most of existing works on fashion recommendation focus on modeling the general compatibility while ignoring the user preferences. In this paper, we present a Personalized Attention Network (PAN) for fashion recommendation. The key component of PAN includes a user encoder, an item encoder and a preference predictor. To modeling users’ diverse interests, we develop an attention network to incorporate the learnt user representation into the item encoder component. More specifically, the attention module consists of a sequential user-aware channel-level and a spatial-level sub-module. Moreover, a novel ranking an user-specific loss, is proposed to capture the interest of different users on the same outfit. To make the training more effective and efficient, a novel user-aware online hard negative mining strategy is proposed. Extensive experiments on Polyvore-U dataset demonstrate the excellence of the proposed system and the effectiveness of different modules.
Huijing Zhan, Jie Lin 0001
ICIP2
2021 Cross-modal discriminant adversarial network
Peng Hu 0002, Xi Peng 0001, Hongyuan Zhu 0002, Jie Lin 0001, Liangli Zhen, Wei Wang 0283, Dezhong Peng
Pattern Recognit.4
2021 Joint Versus Independent Multiview Hashing for Cross-View Retrieval
abstract
Thanks to the low storage cost and high query speed, cross-view hashing (CVH) has been successfully used for similarity search in multimedia retrieval. However, most existing CVH methods use all views to learn a common Hamming space, thus making it difficult to handle the data with increasing views or a large number of views. To overcome these difficulties, we propose a decoupled CVH network (DCHN) approach which consists of a semantic hashing autoencoder module (SHAM) and multiple multiview hashing networks (MHNs). To be specific, SHAM adopts a hashing encoder and decoder to learn a discriminative Hamming space using either a few labels or the number of classes, that is, the so-called flexible inputs. After that, MHN independently projects all samples into the discriminative Hamming space that is treated as an alternative ground truth. In brief, the Hamming space is learned from the semantic space induced from the flexible inputs, which is further used to guide view-specific hashing in an independent fashion. Thanks to such an independent/decoupled paradigm, our method could enjoy high computational efficiency and the capacity of handling the increasing number of views by only using a few labels or the number of classes. For a newly coming view, we only need to add a view-specific network into our model and avoid retraining the entire model using the new and previous views. Extensive experiments are carried out on five widely used multiview databases compared with 15 state-of-the-art approaches. The results show that the proposed independent hashing paradigm is superior to the common joint ones while enjoying high efficiency and the capacity of handling newly coming views.
Peng Hu 0002, Xi Peng 0001, Hongyuan Zhu 0002, Jie Lin 0001, Liangli Zhen, Dezhong Peng
IEEE Trans. Cybern.4
2021 Pose-Normalized and Appearance-Preserved Street-to-Shop Clothing Image Generation and Feature Learning
abstract
We tackle the task of street-to-shop clothing image synthesis. Given a daily person image with a particular clothing item captured in the street scenario, we aim to synthesize the frontal facing view of that item in the shop scenario. This problem has the following challenges: 1) the distinct visual discrepancy between the street and shop scenario; 2) the severe shape deformation of clothing in the presence of an arbitrary human pose; 3) the preservation of fine-grained details during the process of clothing image generation. In this paper, we jointly solve these difficulties by proposing a Pose-Normalized and Appearance-Preserved Generative Adversarial Network (PNAP-GAN). More specifically, conditioned on the clothing-agnostic representation (i.e., clothing landmarks and semantic parsing map), we disentangle the shape and appearance synthesis in a coarse-to-fine framework. Moreover, a semantic embedding loss is introduced to guide the domain transfer in the semantic level (i.e., keeping the clothing attributes). With the synthesized frontal shop image, a pose-normalized representation in complementary to the domain-invariant feature learnt from the original street image are integrated to facilitate the problem of street-to-shop clothing retrieval. Extensive experiments conducted demonstrate the effectiveness of the proposed PNAP-GAN on generating high quality frontal-view images and the excellence of the learnt pose-normalized features on the retrieval task than existing methods. In addition, we demonstrate that the pose-normalized retrieval feature benefits the cross-scenario (i.e., street-to-shop) clothing image generation in a semantic-preserved manner.
Huijing Zhan, Chenyu Yi, Boxin Shi, Jie Lin 0001, Ling-Yu Duan, Alex Chichung Kot
IEEE Trans. Multim.4
2021 Lifelog Image Retrieval Based on Semantic Relevance Mapping
abstract
Lifelog analytics is an emerging research area with technologies embracing the latest advances in machine learning, wearable computing, and data analytics. However, state-of-the-art technologies are still inadequate to distill voluminous multimodal lifelog data into high quality insights. In this article, we propose a novel semantic relevance mapping ( SRM ) method to tackle the problem of lifelog information access. We formulate lifelog image retrieval as a series of mapping processes where a semantic gap exists for relating basic semantic attributes with high-level query topics. The SRM serves both as a formalism to construct a trainable model to bridge the semantic gap and an algorithm to implement the training process on real-world lifelog data. Based on the SRM, we propose a computational framework of lifelog analytics to support various applications of lifelog information access, such as image retrieval, summarization, and insight visualization. Systematic evaluations are performed on three challenging benchmarking tasks to show the effectiveness of our method.
Qianli Xu, Ana Garcia del Molino, Jie Lin 0001, Fen Fang, Vigneshwaran Subbaraju, Liyuan Li, Joo-Hwee Lim
ACM Trans. Multim. Comput. Commun. Appl.3
2021 Multi-GPU Design and Performance Evaluation of Homomorphic Encryption on GPU Clusters
abstract
We present a multi-GPU design, implementation and performance evaluation of the Halevi-Polyakov-Shoup (HPS) variant of the Fan-Vercauteren (FV) levelled Fully Homomorphic Encryption (FHE) scheme. Our design follows a data parallelism approach and uses partitioning methods to distribute the workload in FV primitives evenly across available GPUs. The design is put to address space and runtime requirements of FHE computations. It is also suitable for distributed-memory architectures, and includes efficient GPU-to-GPU data exchange protocols. Moreover, it is user-friendly as user intervention is not required for task decomposition, scheduling or load balancing. We implement and evaluate the performance of our design on two homogeneous and heterogeneous NVIDIA GPU clusters: K80, and a customized P100. We also provide a comparison with a recent shared-memory-based multi-core CPU implementation using two homomorphic circuits as workloads: vector addition and multiplication. Moreover, we use our multi-GPU Levelled-FHE to implement the inference circuit of two Convolutional Neural Networks (CNNs) to perform homomorphically image classification on encrypted images from the MNIST and CIFAR - 10 datasets. Our implementation provides 1 to 3 orders of magnitude speedup compared with the CPU implementation on vector operations. In terms of scalability, our design shows reasonable scalability curves when the GPUs are fully connected.
Ahmad Al Badawi, Bharadwaj Veeravalli, Jie Lin 0001, Xiao Nan, Kazuaki Matsumura, Khin Mi Mi Aung
IEEE Trans. Parallel Distributed Syst.3
2020 Semi-Supervised Multi-Modal Learning with Balanced Spectral Decomposition
abstract
Cross-modal retrieval aims to retrieve the relevant samples across different modalities, of which the key problem is how to model the correlations among different modalities while narrowing the large heterogeneous gap. In this paper, we propose a Semi-supervised Multimodal Learning Network method (SMLN) which correlates different modalities by capturing the intrinsic structure and discriminative correlation of the multimedia data. To be specific, the labeled and unlabeled data are used to construct a similarity matrix which integrates the cross-modal correlation, discrimination, and intra-modal graph information existing in the multimedia data. What is more important is that we propose a novel optimization approach to optimize our loss within a neural network which involves a spectral decomposition problem derived from a ratio trace criterion. Our optimization enjoys two advantages given below. On the one hand, the proposed approach is not limited to our loss, which could be applied to any case that is a neural network with the ratio trace criterion. On the other hand, the proposed optimization is different from existing ones which alternatively maximize the minor eigenvalues, thus overemphasizing the minor eigenvalues and ignore the dominant ones. In contrast, our method will exactly balance all eigenvalues, thus being more competitive to existing methods. Thanks to our loss and optimization strategy, our method could well preserve the discriminative and instinct information into the common space and embrace the scalability in handling large-scale multimedia data. To verify the effectiveness of the proposed method, extensive experiments are carried out on three widely-used multimodal datasets comparing with 13 state-of-the-art approaches.
Peng Hu 0002, Hongyuan Zhu 0002, Xi Peng 0001, Jie Lin 0001
AAAI4
2020 Object Tracking Via ImageNet Classification Scores
abstract
Object tracking is a challenging task in computer vision. The correlation filter based trackers are widely used for visual tracking due to their efficiencies. However, they cannot handle occlusion very well. In this paper, an effective method is proposed for occlusion detection based on high-level classification scores from the Convolutional Neural Network (CNN) trained on the ImageNet dataset. Also, we propose a novel tracking method by holistically considering multiple tracking models trained previously. In each frame, multiple correlation filters are first trained using hierarchical convolutional features, and then progressively selected according to the so-called tracking quality (status). Finally, a linear motion model is adopted to effectively re-detect the lost target. Experimental results have demonstrated that our method achieved good performance for handling occlusion.
Li Wang 0057, Ting Liu 0009, Bing Wang 0003, Jie Lin 0001, Xulei Yang, Gang Wang 0012
ICIP4
2020 Cascaded Mixed-Precision Networks
abstract
There has been a vast literature on Neural Network Compression, either by quantizing network variables to low precision numbers or pruning redundant connections from the network architecture. However, these techniques experience performance degradation when the compression ratio is increased to an extreme extent. In this paper, we propose Cascaded Mixed-precision Networks (CMNs), which are compact yet efficient neural networks without incurring performance drop. CMN is designed as a cascade framework by concatenating a group of neural networks with sequentially increased bitwidth. The execution flow of CMN is conditional on the difficulty of input samples, i.e., easy examples will be correctly classified by going through extremely low-bitwidth networks, and hard examples will be handled by high-bitwidth networks, so that the average compute is reduced. In addition, weight pruning is incorporated into the cascaded framework and jointly optimized with the mixed-precision quantization. To validate this method, we implemented a 2-stage CMN consisting of a binary neural network and a multi-bit (e.g. 8 bits) neural network. Empirical results on CIFAR-100 and ImageNet demonstrate that CMN performs better than state-of-the-art methods, in terms of accuracy and compute.
Xue Geng, Jie Lin 0001
ICIP2
2020 A*3D Dataset: Towards Autonomous Driving in Challenging Environments
abstract
With the increasing global popularity of self-driving cars, there is an immediate need for challenging real-world datasets for benchmarking and training various computer vision tasks such as 3D object detection. Existing datasets either represent simple scenarios or provide only day-time data. In this paper, we introduce a new challenging A*3D dataset which consists of RGB images and LiDAR data with a significant diversity of scene, time, and weather. The dataset consists of high-density images (≈ 10 times more than the pioneering KITTI dataset), heavy occlusions, a large number of nighttime frames (≈ 3 times the nuScenes dataset), addressing the gaps in the existing datasets to push the boundaries of tasks in autonomous driving research to more challenging highly diverse environments. The dataset contains 39K frames, 7 classes, and 230K 3D object annotations. An extensive 3D object detection benchmark evaluation on the A*3D dataset for various attributes such as high density, day-time/night-time, gives interesting insights into the advantages and limitations of training and testing 3D object detection in real-world setting.
Quang-Hieu Pham, Pierre Sevestre, Ramanpreet Singh Pahwa, Huijing Zhan, Chun Ho Pang, Yuda Chen, Armin Mustafa, Vijay Chandrasekhar 0001, Jie Lin 0001
ICRA9
2019 MaxpoolNMS: Getting Rid of NMS Bottlenecks in Two-Stage Object Detectors
abstract
Modern convolutional object detectors have improved the detection accuracy significantly, which in turn inspired the development of dedicated hardware accelerators to achieve real-time performance by exploiting inherent parallelism in the algorithm. Non-maximum suppression (NMS) is an indispensable operation in object detection. In stark contrast to most operations, the commonly-adopted GreedyNMS algorithm does not foster parallelism, which can be a major performance bottleneck. In this paper, we introduce MaxpoolNMS, a parallelizable alternative to the NMS algorithm, which is based on max-pooling classification score maps. By employing a novel multi-scale multi-channel max-pooling strategy, our method is 20x faster than GreedyNMS while simultaneously achieves comparable accuracy, when quantified across various benchmarking datasets, i.e., MS COCO, KITTI and PASCAL VOC. Furthermore, our method is better suited for hardware-based acceleration than GreedyNMS.
Lile Cai, Zhe Wang 0019, Jie Lin 0001, Chuan-Sheng Foo, Mohamed M. Sabry, Vijay Chandrasekhar 0001
CVPR4
2019 Dataflow-Based Joint Quantization for Deep Neural Networks
abstract
This paper addresses a challenging problem - how to reduce energy consumption without incurring performance drop when deploying deep neural networks (DNNs) at the inference stage. In order to alleviate the computation and storage burdens, we propose a novel dataflow-based joint quantization approach with the hypothesis that a fewer number of quantization operations would incur less information loss and thus improve the final performance. It first introduces a quantization scheme with efficient bit-shifting and rounding operations to represent network parameters and activations in low precision. Then it re-structures the network architectures to form unified modules for optimization on the quantized model. Extensive experiments on ImageNet and KITTI validate the effectiveness of our model, demonstrating that state-of-the-art results for various tasks can be achieved by this quantized model. Besides, we designed and synthesized an RTL model to measure the hardware costs among various quantization methods. For each quantization operation, it reduces area cost by about 15 times and energy consumption by about 9 times, compared to a strong baseline.
Xue Geng, Jie Fu 0001, Jie Lin 0001, Mohamed M. Sabry, Christopher Joseph Pal, Vijay Chandrasekhar 0001
DCC4
2019 Beyond Ranking Loss: Deep Holographic Networks for Multi-Label Video Search
abstract
In this paper, we propose Deep Holographic Networks (DHN) to learn similarity metrics of videos for multi-label video search. DHN introduces a holographic composition layer to explicitly encode similarity metrics at intermediate layer of the network, instead of conventional deep metric learning approaches driven by ranking losses. The holographic composition layer is parameter-free and enables less memory footprint compared with state-of-the-art. Towards multi-label video search at large scale, we present a new video benchmark built upon the YouTube-8M dataset. Extensive evaluations on this dataset demonstrate that DHN performs better than traditional deep metric learning approaches as well as other compositional networks.
Zhuo Chen 0006, Jie Lin 0001, Zhe Wang 0019, Vijay Chandrasekhar 0001, Weisi Lin
ICIP2
2019 Learning Hierarchical Features for Visual Object Tracking With Recursive Neural Networks
abstract
Recently, deep learning has achieved very promising results in visual object tracking. Deep neural networks in existing tracking methods require a lot of training data to learn a large number of parameters. However, training data is not sufficient for visual object tracking as annotations of a target object are only available in the first frame of a test sequence. In this paper, we propose to learn hierarchical features for visual object tracking by using tree structure based Recursive Neural Networks (RNN), which have a relatively small number of parameters compared to other deep neural networks (e.g. Convolutional Neural Networks (CNN)) due to all basic modules in RNN share only one set of parameters. Experimental results demonstrate that our feature learning algorithm can significantly improve tracking performance on benchmark datasets.
Li Wang 0057, Ting Liu 0009, Bing Wang 0003, Jie Lin 0001, Xulei Yang, Gang Wang 0012
ICIP4
2019 Optimizing the Bit Allocation for Compression of Weights and Activations of Deep Neural Networks
abstract
For most real-time implementations of deep artificial neural networks, both weights and intermediate layer activations have to be stored and loaded for processing. Compression of both is often advisable to mitigate the memory bottleneck. In this paper, we propose a bit allocation framework for compressing the weights and activations of deep neural networks. The differentiability of input-output relationships for all network layers allows us to relate the neural network output accuracy to the bit-rate of quantized weights and layer activations. We formulate a Lagrangian optimization framework that finds the optimum joint bit allocation among all intermediate activation layers and weights. Our method obtains excellent results on two deep neural networks, VGG-16 and ResNet-50. Without requiring re-training, it outperforms other state-of-the-art neural network compression methods.
Wang Zhe, Jie Lin 0001, Vijay Chandrasekhar 0001, Bernd Girod
ICIP2
2019 TEA-DNN: the Quest for Time-Energy-Accuracy Co-optimized Deep Neural Networks
abstract
Embedded deep learning platforms have witnessed two simultaneous improvements. First, the accuracy of convolutional neural networks (CNNs) has been significantly improved through the use of automated neural-architecture search (NAS) algorithms to determine CNN structure. Second, there has been increasing interest in developing hardware accelerators for CNNs that provide improved inference performance and energy consumption compared to GPUs. Such embedded deep learning platforms differ in the amount of compute resources and memory-access bandwidth, which would affect performance and energy consumption of CNNs. It is therefore critical to consider the available hardware resources in the network architecture search. To this end, we introduce TEA-DNN, a NAS algorithm targeting multi-objective optimization of execution time, energy consumption, and classification accuracy of CNN workloads on embedded architectures. TEA-DNN leverages energy and execution time measurements on embedded hardware when exploring the Pareto-optimal curves across accuracy, execution time, and energy consumption and does not require additional effort to model the underlying hardware. We apply TEA-DNN for image classification on actual embedded platforms (NVIDIA Jetson TX2 and Intel Movidius Neural Compute Stick). We highlight the Pareto-optimal operating points that emphasize the necessity to explicitly consider hardware characteristics in the search process. To the best of our knowledge, this is the most comprehensive study of Pareto-optimal models across a range of hardware platforms using actual measurements on hardware to obtain objective values.
Lile Cai, Anne-Maelle Barneche, Arthur Herbout, Chuan-Sheng Foo, Jie Lin 0001, Vijay Chandrasekhar 0001, Mohamed M. Sabry
ISLPED5
2019 Trademark image retrieval via transformation-invariant deep hashing
Zhaoqiang Xia, Jie Lin 0001, Xiaoyi Feng
J. Vis. Commun. Image Represent.2
2019 Codebook-Free Compact Descriptor for Scalable Visual Search
abstract
The MPEG compact descriptors for visual search (CDVS) is a standard toward image matching and retrieval. To achieve high retrieval accuracy over a large scale image/video dataset, recent research efforts have demonstrated that employing extremely high-dimensional descriptors such as the Fisher vector (FV) and the vector of locally aggregated descriptors (VLAD) can yield good performance. Since the FV (or VLAD) possesses high discriminability but small visual vocabulary, it has been adopted by CDVS to construct a global compact descriptor. In this paper, we study the development of global compact descriptors in the completed CDVS standard and the emerging compact descriptors for video analysis (CDVA) standard, in which we formulate the FV (or VLAD) compression as a resource-constrained optimization problem. Accordingly, we propose a codebook-free aggregation method via dual selection to generate a global compact visual descriptor, which supports fast and accurate feature matching free of large visual codebooks, fulfilling the low memory requirement of mobile visual search at significantly reduced latency. Specifically, we investigate both sample-specific Gaussian component redundancy and bit dependency within a binary aggregated descriptor to produce compact binary codes. Our technique contributes to the scalable compressed Fisher vector (SCFV) adopted by the CDVS standard. Moreover, the SCFV descriptor is currently serving as the frame-level hand-crafted video feature, which inspires the inheritance of CDVS descriptors for the emerging CDVA standard. Furthermore, we investigate the positive complementary effect of our standard compliant compact descriptor and deep learning based features extracted from convolutional neural networks with significant mean average precision gains. Extensive evaluation over benchmark databases shows the significant merits of the codebook-free binary codes for scalable visual search.
Yuwei Wu 0001, Feng Gao 0014, Jie Lin 0001, Vijay Chandrasekhar 0001, Junsong Yuan 0001, Ling-Yu Duan
IEEE Trans. Multim.4
2018 Hardware-Aware Softmax Approximation for Deep Neural Networks
Xue Geng, Jie Lin 0001, Anmin Kong, Mohamed M. Sabry, Vijay Chandrasekhar 0001
ACCV (4)2
2018 Gated Square-Root Pooling for Image Instance Retrieval
abstract
Recently Convolutional Neural Networks (CNNs) have achieved great success in different fields including image instance retrieval. However traditional global pooling approaches fail to capture all possible discriminative information of CNN activations and treat activations over channels equally regardless of the different importance between channels. In this work, we focus on the mentioned problem of global feature pooling over CNN activations for image instance retrieval. We make two contributions. First, we introduce a channel-wise SQUare-root (SQU) pooling (2-norm) approach, which makes better use of information over activation maps and is superior to Average (1-norm) and Max pooling (infinity norm), in the context of instance retrieval. Second, we further improve SQU by learning a gating function that weights the contributions of different channels, in an end-to-end manner. Extensive experiments on 6 benchmark datasets show that the proposed strategies achieve considerable improvements over state-of-the-art.
Ziqian Chen, Jie Lin 0001, Vijay Chandrasekhar 0001, Ling-Yu Duan
ICIP2
2017 Compression of Deep Neural Networks for Image Instance Retrieval
abstract
Image instance retrieval is the problem of retrieving images from a database which contain the same object. Convolutional Neural Network (CNN) based descriptors are becoming the dominant approach for generating global image descriptors for the instance retrieval problem. One major drawback of CNN-based global descriptors is that uncompressed deep neural network models require hundreds of megabytes of storage making them inconvenient to deploy in mobile applications or in custom hardware. In this work, we study the problem of neural network model compression focusing on the image instance retrieval task. We study quantization, coding, pruning and weight sharing techniques for reducing model size for the instance retrieval problem. We provide extensive experimental results on the trade-off between retrieval performance and model size for different types of networks on several data sets providing the most comprehensive study on this topic. We compress models to the order of a few MBs: two orders of magnitude smaller than the uncompressed models while achieving negligible loss in retrieval performance1.
Vijay Chandrasekhar 0001, Jie Lin 0001, Qianli Liao, Olivier Morère, Antoine Veillard, Ling-Yu Duan, Tomaso A. Poggio
DCC2
2017 Compact Deep Invariant Descriptors for Video Retrieval
abstract
With emerging demand for large-scale video analysis, the Motion Picture Experts Group (MPEG) initiated the Compact Descriptor for Video Analysis (CDVA) standardization in 2014. In this work, we develop novel deep-learning features and incorporate them into the well-established CDVA evaluation framework to study its effectiveness in video analysis. In particular, we propose a Nested Invariance Pooling (NIP) method to obtain compact and robust Convolutional Neural Network (CNNs) descriptors. The CNNs descriptors are generated by applying three different pooling operations to the feature maps of CNNs in a nested way towards rotation and scale invariant feature representation. In particular, the rational, advantages and performance on the combination of CNNs and handcrafted descriptors are provided to better investigate the complementary effects of deep learnt and handcrafted features. Extensive experimental results show that the proposed CNNs descriptors outperform both state-of-the-art CNNs descriptors and canonical handcrafted descriptors adopted in CDVA Experimental Model (CXM) with significant mAP gains of 11.3% and 4.7%, respectively. Moreover, the combination of NIP derived deep invariant descriptors and handcrafted descriptors not only fulfills the lowest bitrate budget of CDVA, but also significantly advances the performance of CDVA core techniques.
Yihang Lou, Jie Lin 0001, Shiqi Wang 0001, Jie Chen 0006, Vijay Chandrasekhar 0001, Ling-Yu Duan, Tiejun Huang 0001, Alex Chichung Kot, Wen Gao 0001
DCC3
2017 Deep regional feature pooling for video matching
abstract
In this work, we study the problem of deep global descriptors for video matching with regional feature pooling. We aim to analyze the joint effect of ROI (Region of Interest) size and pooling moment on video matching performance. To this end, we propose to mathematically model the distribution of video matching function with a pooling function nested in. Matching performance can be estimated by the separability of these class-conditional distributions between matching and non-matching pairs. Empirical studies on the challenging MPEG CDVA dataset demonstrate that performance trends are consistent with the estimation and experimental results, though the theoretical model is largely simplified compared to video matching and retrieval in practice.
Jie Lin 0001, Vijay Chandrasekhar 0001, Yihang Lou, Shiqi Wang 0001, Ling-Yu Duan, Tiejun Huang 0001, Alex Chichung Kot
ICIP2
2017 A Multi-Block N-ary trie structure for exact r-neighbour search in hamming space
abstract
This paper proposes a novel algorithm to solve the exact r-neighbour search problem in Hamming space. Existing r-neighbour search methods typically adopt hash table to index binary codes. Given a query, existing approaches search the nearest neighbours by checking all buckets of a Hamming ball centered at the query. The problem is these methods spend most of search time visiting empty buckets (lookup misses). In this paper, we adopt trie structure to index binary codes. We consider several continuous bits of a binary string as a block and use it as an atomic indexing element in trie structure, which is efficient in access speed and memory usage. Our method searches the nearest neighbours of a query by utilizing the records of nodes in trie structure to avoid lookup misses. We name the proposed indexing structure as Multi-Block N-ary Trie (MBNT). A theoretical analysis is given to prove that MBNT has less time cost than other hash-based methods. Extensive results show that MBNT outperforms state-of-the-art algorithms on several large scale benchmarks.
Ling-Yu Duan, Zhe Wang 0019, Jie Lin 0001, Vijay Chandrasekhar 0001, Tiejun Huang 0001
ICIP4
2017 Region average pooling for context-aware object detection
abstract
Object detection has been a key task in computer vision with deep convolutional neural networks being a significant performer. We propose a method named Region Average Pooling that leverages object co-occurrence to improve object detection performance. Given regions of interest in an image, our method augments object detection networks with pooled contextual features from other regions of interest in the scene. We implement our scheme and evaluate it on the Pascal Visual Object Classes (VOC) 2007 and Microsoft Common Objects in Context (MS COCO) datasets. When used as part of the Faster R-CNN object detection framework with VGG-16, we show an increase in mAP from 24.2% to 25.5% over baseline Faster R-CNN and Global Average Pooling when testing on MS COCO.
Kingsley Kuan, Gaurav Manek, Jie Lin 0001, Yuan Fang 0001, Vijay Chandrasekhar 0001
ICIP3
2017 Object Detection Meets Knowledge Graphs
abstract
Object detection in images is a crucial task in computer vision, with important applications ranging from security surveillance to autonomous vehicles. Existing state-of-the-art algorithms, including deep neural networks, only focus on utilizing features within an image itself, largely neglecting the vast amount of background knowledge about the real world. In this paper, we propose a novel framework of knowledge-aware object detection, which enables the integration of external knowledge such as knowledge graphs into any object detection algorithm. The framework employs the notion of semantic consistency to quantify and generalize knowledge, which improves object detection through a re-optimization process to achieve better consistency with background knowledge. Finally, empirical evaluation on two benchmark datasets show that our approach can significantly increase recall by up to 6.3 points without compromising mean average precision, when compared to the state-of-the-art baseline.
Yuan Fang 0001, Kingsley Kuan, Jie Lin 0001, Cheston Tan, Vijay Chandrasekhar 0001
IJCAI3
2017 DeepHash for Image Instance Retrieval: Getting Regularization, Depth and Fine-Tuning Right
abstract
This work focuses on representing very high-dimensional global image descriptors using very compact 64-1024 bit binary hashes for instance retrieval. We propose DeepHash: a hashing scheme based on deep networks. Key to making DeepHash work at extremely low bitrates are three important considerations -- regularization, depth and fine-tuning -- each requiring solutions specific to the hashing problem. In-depth evaluation shows that our scheme outperforms state-of-the-art methods over several benchmark datasets for both Fisher Vectors and Deep Convolutional Neural Network features, by up to 8.5% over other schemes. The retrieval performance with 256-bit hashes is close to that of the uncompressed floating point features -- a remarkable 512x compression.
Jie Lin 0001, Olivier Morère, Antoine Veillard, Ling-Yu Duan, Hanlin Goh, Vijay Chandrasekhar 0001
ICMR1
2017 Nested Invariance Pooling and RBM Hashing for Image Instance Retrieval
abstract
The goal of this work is the computation of very compact binary hashes for image instance retrieval. Our approach has two novel contributions. The first one is Nested Invariance Pooling (NIP), a method inspired from i-theory, a mathematical theory for computing group invariant transformations with feed-forward neural networks. NIP is able to produce compact and well-performing descriptors with visual representations extracted from convolutional neural networks. We specifically incorporate scale, translation and rotation invariances but the scheme can be extended to any arbitrary sets of transformations. We also show that using moments of increasing order throughout nesting is important. The NIP descriptors are then hashed to the target code size (32-256 bits) with a Restricted Boltzmann Machine with a novel batch-level regularization scheme specifically designed for the purpose of hashing (RBMH). A thorough empirical evaluation with state-of-the-art shows that the results obtained both with the NIP descriptors and the NIP+RBMH hashes are consistently outstanding across a wide range of datasets.
Olivier Morère, Jie Lin 0001, Antoine Veillard, Ling-Yu Duan, Vijay Chandrasekhar 0001, Tomaso A. Poggio
ICMR2
2017 Low Bit-rate 3D feature descriptors for depth data from Kinect-style sensors
Sai Manoj Prakhya, Weisi Lin, Vijay Chandrasekhar 0001, Jie Lin 0001
Signal Process. Image Commun.5
2017 Deep convolutional hashing using pairwise multi-label supervision for large-scale visual search
Zhaoqiang Xia, Xiaoyi Feng, Jie Lin 0001, Abdenour Hadid
Signal Process. Image Commun.3
2017 Towards Detection of Bus Driver Fatigue Based on Robust Visual Analysis of Eye State
abstract
Driver's fatigue is one of the major causes of traffic accidents, particularly for drivers of large vehicles (such as buses and heavy trucks) due to prolonged driving periods and boredom in working conditions. In this paper, we propose a vision-based fatigue detection system for bus driver monitoring, which is easy and flexible for deployment in buses and large vehicles. The system consists of modules of head-shoulder detection, face detection, eye detection, eye openness estimation, fusion, drowsiness measure percentage of eyelid closure (PERCLOS) estimation, and fatigue level classification. The core innovative techniques are as follows: 1) an approach to estimate the continuous level of eye openness based on spectral regression; and 2) a fusion algorithm to estimate the eye state based on adaptive integration on the multimodel detections of both eyes. A robust measure of PERCLOS on the continuous level of eye openness is defined, and the driver states are classified on it. In experiments, systematic evaluations and analysis of proposed algorithms, as well as comparison with ground truth on PERCLOS measurements, are performed. The experimental results show the advantages of the system on accuracy and robustness for the challenging situations when a camera of an oblique viewing angle to the driver's face is used for driving state monitoring.
Bappaditya Mandal, Liyuan Li, Gang S. Wang, Jie Lin 0001
IEEE Trans. Intell. Transp. Syst.4
2017 HNIP: Compact Deep Invariant Representations for Video Matching, Localization, and Retrieval
abstract
With emerging demand for large-scale video analysis, MPEG initiated the compact descriptor for video analysis (CDVA) standardization in 2014. Beyond handcrafted descriptors adopted by the current MPEG-CDVA reference model, we study the problem of deep learned global descriptors for video matching, localization, and retrieval. First, inspired by a recent invariance theory, we propose a nested invariance pooling (NIP) method to derive compact deep global descriptors from convolutional neural networks (CNNs), by progressively encoding translation, scale, and rotation invariances into the pooled descriptors. Second, our empirical studies have shown that a sequence of well designed pooling moments (e.g., max or average) may drastically impact video matching performance, which motivates us to design hybrid pooling operations via NIP (HNIP). HNIP has further improved the discriminability of deep global descriptors. Third, the technical merits and performance improvements by combining deep and handcrafted descriptors are provided to better investigate the complementary effects. We evaluate the effectiveness of HNIP within the well-established MPEG-CDVA evaluation framework. The extensive experiments have demonstrated that HNIP outperforms the state-of-the-art deep and canonical handcrafted descriptors with significant mAP gains of 5.5% and 4.7%, respectively. In particular the combination of HNIP incorporated and handcrafted global descriptors has significantly boosted the performance of CDVA core techniques with comparable descriptor size.
Jie Lin 0001, Ling-Yu Duan, Shiqi Wang 0001, Yihang Lou, Vijay Chandrasekhar 0001, Tiejun Huang 0001, Alex Chichung Kot, Wen Gao 0001
IEEE Trans. Multim.1
2016 Tiny Descriptors for Image Retrieval with Unsupervised Triplet Hashing
abstract
A typical image retrieval pipeline starts with the comparison of global descriptors from a large database to find a short list of candidate matches. A good image descriptor is key to the retrieval pipeline and should reconcile two contradictory requirements: providing recall rates as high as possible and being as compact as possible for fast matching. Following the recent successes of Deep Convolutional Neural Networks (DCNN) for large scale image classification, descriptors extracted from DCNNs are increasingly used in place of the traditional hand crafted descriptors such as Fisher Vectors (FV) with better retrieval performances. Nevertheless, the dimensionality of a typical DCNN descriptor-extracted either from the visual feature pyramid or the fully-connected layers-remains quite high at several thousands of scalar values. In this paper, we propose Unsupervised Triplet Hashing (UTH), a fully unsupervised method to compute extremely compact binary hashes-in the 32-256 bits range-from high-dimensional global descriptors. UTH consists of two successive deep learning steps. First, Stacked Restricted Boltzmann Machines (SRBM), a type of unsupervised deep neural nets, are used to learn binary embedding functions able to bring the descriptor size down to the desired bitrate. SRBMs are typically able to ensure a very high compression rate at the expense of loosing some desirable metric properties of the original DCNN descriptor space. Then, triplet networks, a rank learning scheme based on weight sharing nets is used to fine-tune the binary embedding functions to retain as much as possible of the useful metric properties of the original space. A thorough empirical evaluation conducted on multiple publicly available dataset using DCNN descriptors shows that our method is able to significantly outperform state-of-the-art unsupervised schemes in the target bit range.
Jie Lin 0001, Olivier Morère, Julie Petta, Vijay Chandrasekhar 0001, Antoine Veillard
DCC1
2016 Egocentric activity recognition with multimodal fisher vector
abstract
With the increasing availability of wearable devices, research on egocentric activity recognition has received much attention recently. In this paper, we build a Multimodal Egocentric Activity dataset which includes egocentric videos and sensor data of 20 fine-grained and diverse activity categories. We present a novel strategy to extract temporal trajectory-like features from sensor data. We propose to apply the Fisher Kernel framework to fuse video and temporal enhanced sensor features. Experiment results show that with careful design of feature extraction and fusion algorithm, sensor data can enhance information-rich video data. We make publicly available the Multimodal Egocentric Activity dataset to facilitate future research.
Sibo Song, Ngai-Man Cheung, Vijay Chandrasekhar 0001, Bappaditya Mandal, Jie Lin 0001
ICASSP5
2016 A practical guide to CNNs and Fisher Vectors for image instance retrieval
Vijay Chandrasekhar 0001, Jie Lin 0001, Olivier Morère, Hanlin Goh, Antoine Veillard
Signal Process.2
2016 Overview of the MPEG-CDVS Standard
abstract
Compact descriptors for visual search (CDVS) is a recently completed standard from the ISO/IEC moving pictures experts group (MPEG). The primary goal of this standard is to provide a standardized bitstream syntax to enable interoperability in the context of image retrieval applications. Over the course of the standardization process, remarkable improvements were achieved in reducing the size of image feature data and in reducing the computation and memory footprint in the feature extraction process. This paper provides an overview of the technical features of the MPEG-CDVS standard and summarizes its evolution.
Ling-Yu Duan, Vijay Chandrasekhar 0001, Jie Chen 0006, Jie Lin 0001, Zhe Wang 0019, Tiejun Huang 0001, Bernd Girod, Wen Gao 0001
IEEE Trans. Image Process.4
2015 Compact Global Descriptors for Visual Search
abstract
The first step in an image retrieval pipeline consists of comparing global descriptors from a large database to find a short list of candidate matching images. The more compact the global descriptor, the faster the descriptors can be compared for matching. State-of-the-art global descriptors based on Fisher Vectors are represented with tens of thousands of floating point numbers. While there is significant work on compression of local descriptors, there is relatively little work on compression of high dimensional Fisher Vectors. We study the problem of global descriptor compression in the context of image retrieval, focusing on extremely compact binary representations: 64-1024 bits. Motivated by the remarkable success of deep neural networks in recent literature, we propose a compression scheme based on deeply stacked Restricted Boltzmann Machines (SRBM), which learn lower dimensional non-linear subspaces on which the data lie. We provide a thorough evaluation of several state-of-the-art compression schemes based on PCA, Locality Sensitive Hashing, Product Quantization and greedy bit selection, and show that the proposed compression scheme outperforms all existing schemes.
Vijay Chandrasekhar 0001, Jie Lin 0001, Olivier Morère, Antoine Veillard, Hanlin Goh
DCC2
2015 Optimizing Binary Fisher Codes for Visual Search
abstract
Fisher vectors (FV) aggregated from local invariant features (e.g., SIFT) is one of the state-of-the-art descriptors for visual search, due to high discriminability but small visual vocabulary. Nevertheless, a high-dimensional FV needs to be compressed into a compact descriptor for light storage and high matching eficiency. In this paper, we formulate the FV compression as a resource-constrained optimization problem. Our goal is to maximize search performance subject to the constraints of descriptor compactness, compression complexity in terms of memory usage and time cost. Accordingly, we present a selective binary Fisher codes (SBFC) to compress the raw FV. Firstly, to fulfill the constraint of compression complexity, we binarize the FV by a sign function, Secondly, we propose to select discriminative bits from the binarized FV (BFC) to maximize search performance, subject to the constraint of descriptor compactness. Extensive experiments over MPEG Compact Descriptor for Visual Search (CDVS) benchmark datasets have shown that S-BFC significantly improves search performance at a smaller descriptor size as well as much lower complexity, compared with the state-of-the-art FV compression algorithms like Hashing and Product Quantziation (PQ). A simplified version of SBFC, SBFC LS has been adopted by the MPEG CDVS standard. In the CDVS evaluation framework, SBFC LS has achieved promising performance mean Average Precision (mAP) 83% on average at much lower memory cost of 40KB.
Zhe Wang 0019, Ling-Yu Duan, Jie Lin 0001, Jie Chen 0006, Tiejun Huang 0001, Wen Gao 0001
DCC3
2015 An efficient coding framework for compact descriptors extracted from video sequence
abstract
Towards effective and efficient image matching or retrieval tasks, the emerging MPEG standard, named Compact Descriptors for Visual Search (CDVS), has fulfilled compact descriptors for still images, consisting of compressed local and global descriptor. Nevertheless, the frame-level coding of CDVS descriptors from a video sequence does not address the inter-frame redundancy issue, which may consume considerable bandwidth and storage resources. In this work, we propose an efficient coding framework of CDVS descriptors to generate compact descriptors for video sequences. For local descriptors, we propose a multiple reference predictive technique to exploit the temporal correlation of local descriptors and location coordinates over a sequence of frames. To further improve the prediction performance, keypoint tracking is applied to identify temporally repeated keypoints. For global descriptors, a propagation coding way is employed to compress the global descriptors of adjacent frames. The empirical evaluation has shown that the proposed coding approach has yielded a low bit rate of less than 40kbps on average, while maintaining comparable matching and retrieval performance. Compared to the sequence of original frame-level CDVS descriptors, the proposed approach has achieved over 25× bit rate reduction.
Zhangshuai Huang, Ling-Yu Duan, Jie Lin 0001, Shiqi Wang 0001, Siwei Ma 0001, Tiejun Huang 0001
ICIP3
2015 Co-regularized deep representations for video summarization
abstract
Compact keyframe-based video summaries are a popular way of generating viewership on video sharing platforms. Yet, creating relevant and compelling summaries for arbitrarily long videos with a small number of keyframes is a challenging task. We propose a comprehensive keyframe-based summarization framework combining deep convolutional neural networks and restricted Boltzmann machines. An original co-regularization scheme is used to discover meaningful subject-scene associations. The resulting multimodal representations are then used to select highly-relevant keyframes. A comprehensive user study is conducted comparing our proposed method to a variety of schemes, including the summarization currently in use by one of the most popular video sharing websites. The results show that our method consistently outperforms the baseline schemes for any given amount of keyframes both in terms of attractiveness and in-formativeness. The lead is even more significant for smaller summaries.
Olivier Morère, Hanlin Goh, Antoine Veillard, Vijay Chandrasekhar 0001, Jie Lin 0001
ICIP5
2015 Hierarchical multi-VLAD for image retrieval
abstract
Constructing discriminative feature descriptors is crucial towards effective image retrieval. The state-of-the-art powerful global descriptor for this purpose is Vector of Locally Aggregated Descriptors (VLAD). Given a set of local features (say, SIFT) extracted from an image, the VLAD is generated by quantizing local features with a small visual vocabulary (64 to 512 centroids), aggregating the residual statistics of quantized features for each centroid and concatenating the aggregated residual vectors from each centroid. One can increase the search accuracy by increasing the size of vocabulary (from hundreds to hundreds of thousands), which, however, it leads to heavy computation cost with flat quantization. In this paper, we propose a hierarchical multi-VLAD to seek the tradeoff between descriptor discriminability and computation complexity. We build up a tree-structured hierarchical quantization (TSHQ) to accelerate the VLAD computation with a large vocabulary. As quantization error may propagate from root to leaf node (centroid) with TSHQ, we introduce multi-VLAD, which constructing a VLAD descriptor for each level of the vocabulary tree, so as to compensate for the quantization error at that level. Extensive evaluation over benchmark datasets has shown that the proposed approach outperforms state-of-the-art in terms of retrieval accuracy, fast extraction, as well as light memory cost.
Ling-Yu Duan, Jie Lin 0001, Zhe Wang 0019, Tiejun Huang 0001
ICIP3
2015 Hamming Compatible Quantization for Hashing
Zhe Wang 0019, Ling-Yu Duan, Jie Lin 0001, Tiejun Huang 0001, Wen Gao 0001
IJCAI3
2015 Weighted Component Hashing of Binary Aggregated Descriptors for Fast Visual Search
abstract
Towards low bit rate mobile visual search, recent works have proposed to aggregate the local features and compress the aggregated descriptor (such as Fisher vector, the vector of locally aggregated descriptors) for low latency query delivery as well as moderate search complexity. Even though Hamming distance can be computed very fast, the computational cost of exhaustive linear search over the binary descriptors grows linearly with either the length of a binary descriptor or the number of database images. In this paper, we propose a novel weighted component hashing (WeCoHash) algorithm for long binary aggregated descriptors to significantly improve search efficiency over a large scale image database. Accordingly, the proposed WeCoHash has attempted to address two essential issues in Hashing algorithms: “what to hash” and “how to search.” “What to hash” is tackled by a hybrid approach, which utilizes both image-specific component (i.e., visual word) redundancy and bit dependency within each component of a binary aggregated descriptor to produce discriminative hash values for bucketing. “How to search” is tackled by an adaptive relevance weighting based on the statistics of hash values. Extensive comparison results have shown that WeCoHash is at least 20 times faster than linear search and 10 times faster than local sensitive hash (LSH) when maintaining comparable search accuracy. In particular , the WeCoHash solution has been adopted by the emerging MPEG compact descriptor for visual search (CDVS) standard to significantly speed up the exhaustive search of the binary aggregated descriptors.
Ling-Yu Duan, Jie Lin 0001, Zhe Wang 0019, Tiejun Huang 0001, Wen Gao 0001
IEEE Trans. Multim.2
2014 Joint optimization of JPEG quantization table and coefficient thresholding for low bitrate mobile visual search
abstract
Low latency query delivery over wireless network is a key problem for mobile visual search. Extracting compact descriptors directly on the mobile device is computational expensive, an alternate approach is to send highly compressed JPEG query images. As JPEG baseline optimizes the rate-distortion from a perceptual perspective rather than maintaining search performance, recent work proposed to learn a feature-preserving JPEG quantization table for improved search accuracy. However, this method is data-dependent and the quantization table cannot adapt to image blocks. To address these issues, we propose to jointly optimize the JPEG quantization table and coefficient thresholding. The matching score between uncompressed image and its compressed JPEG image is employed as the distortion measure to avoid time consuming image labeling, and coefficient thresholding eliminates the redundant coefficients. Extensive experiments on benchmark datasets show that our approach obtains superior performance than state-of-the-art at low bitrates, meanwhile, it consumes lower cost including processing time, memory and battery on mobile device.
Ling-Yu Duan, Jie Lin 0001, Tiejun Huang 0001, Wen Gao 0001
ICIP3
2014 Component hashing of variable-length binary aggregated descriptors for fast image search
abstract
Compact locally aggregated binary features have shown great advantages in image search. As the exhaustive linear search in Hamming space still entails too much computational complexity for large datasets, recent works proposed to directly use binary codes as hash indices, yielding a dramatic increase in speedup. However, these methods cannot be directly applied to variable-length binary features. In this paper, we propose a Component Hashing (CoHash) algorithm to handle the variable-length binary aggregated descriptors indexing for fast image search. The main idea is to decompose the distance measure between variable-length descriptors into aligned component-to-component matching problems independently, and build multiple hash tables for the visual word components. Given a query, its candidate neighbors are found by using the query binary sub-vectors as indices into their corresponding hash tables. In particular, a bit selection based on conditional mutual information maximization is proposed to reduce the dimensionality of visual word components, which provides a light storage of indices and balances the retrieval accuracy and search cost. Extensive experiments on benchmark datasets show that our approach is 20~25 times faster than linear search, without any noticeable retrieval performance loss.
Zhe Wang 0019, Ling-Yu Duan, Jie Lin 0001, Tiejun Huang 0001, Wen Gao 0001, Miroslaw Bober
ICIP3
2013 On the interoperability of local descriptors compression
abstract
There are a number of component technologies that are useful for visual search, including format of visual descriptors, descriptor extraction process, as well as indexing, and matching algorithms. As a minimum, the format of descriptors as well as parts of their extraction process should be defined to ensure interoperability. In this paper, we study the problem of interoperability among compressed local descriptors at different bit-rates; that is, allowing effective and efficient comparison of compact descriptors, which is fundamentally important to mobile visual search applications. We propose to combine feature transform and multi-stage vector quantization to implement the interoperability of compact local descriptors. First, an orthogonal transform (e.g. Principle component analysis, PCA) is employed to eliminate the correlation between local feature dimensions, which improves the performance of compressed domain descriptor matching with the well-aligned distance computing of sorted important features in transform space. Second, a multi-stage vector quantization (MSVQ) is applied to generate compact codes for local descriptors. At light quantization tables, MSVQ takes advantage of the transform domain features to properly allocate different budgets to each group of transformed feature dimensions, respectively. The interoperability between compressed descriptors at different bit rates can be achieved by the descriptors' fast matching in the orthogonal feature space. In other words, descriptor decoding into the original feature space (SIFT space) is unnecessary, as the distance can be calculated by pre-computed lookup tables. In particular, such efficient matching in transform domain is significant for large-scale visual search. Over a set of benchmark datasets, we have reported superior performance over state-of-the-arts.
Jie Chen 0006, Ling-Yu Duan, Jie Lin 0001, Rongrong Ji, Tiejun Huang 0001, Wen Gao 0001
ICASSP3
2013 Robust fisher codes for large scale image retrieval
abstract
Fisher vectors (FV) have shown great advantages in large scale visual search. However, traditional FV suffers from noisy local descriptors, which may deteriorate the FV discriminative power. In this paper, we propose a robust Fisher vectors (RFV). To fulfill fast search and light storage over a large scale image dataset, we employ a simple binarization method to compress RFV to generate compact robust Fisher codes (RFC). Extensive comparison experiments on benchmark datasets have shown that both RFV and RFC outperforms the state-of-the-art performance. The scalability of RFC has been validated on a dataset of over 1 million images as well.
Jie Lin 0001, Ling-Yu Duan, Tiejun Huang 0001, Wen Gao 0001
ICASSP1
2013 A novel pair-wise image matching strategy with compact descriptors
abstract
In this paper, we address the problem of pair-wise image matching which determines whether two images depict the same objects or scenes. SIFT-like local descriptor-based matching is the most widely adopted method for this purpose and has achieved the state-of-the-art performance. However, local descriptor-based methods usually fail when an image pair contains multiple similar local regions. This problem becomes more serious when coming to limited computational and storage resources. Although global descriptors, e.g., Fisher Vectors, can solve this issue, it is difficult for global descriptors to distinguish images containing different objects of the same class. Therefore, we propose a novel strategy to integrate local and global descriptors for better matching accuracy. To further fulfill the efficiency requirement of applications, we combine dimension reduction and product quantization to obtain compact descriptors and speed up the matching process with pre-computed lookup tables. Extensive comparisons to the state-of-the-art methods demonstrate our advantages in both matching accuracy and efficiency.
Ling-Yu Duan, Jie Lin 0001, Tiejun Huang 0001
ICIP3
2013 Compact descriptors for mobile visual search and MPEG CDVS standardization
abstract
In this paper, we present the state-of-the-art compact descriptors for mobile visual search. In particular, we introduce our MPEG contributions in global descriptor aggregation and local descriptor compression, which have been adopted by the ongoing MPEG standardization of compact descriptor for visual search (CDVS). Standardization progress will be introduced. Other issues including visual object databases and MPEG CDVS impact on visual search industry will be discussed as well.
Ling-Yu Duan, Feng Gao 0014, Jie Chen 0006, Jie Lin 0001, Tiejun Huang 0001
ISCAS4
2012 Learning multiple codebooks for low bit rate mobile visual search
abstract
Compressing a query image's signature via vocabulary coding is an effective approach to low bit rate mobile visual search. State-of-the-art methods concentrate on offline learning a codebook from an initial large vocabulary. Over a large heterogeneous reference database, learning a single codebook may not suffice for maximally removing redundant codewords for vocabulary based compact descriptor. In this paper, we propose to learn multiple codebooks (m-Codebooks) for extremely compressing image signatures. A query-specific codebook (q-Codebook) is online generated at both client and server sides by adaptively weighting the off-line learned multiple codebooks. The q-Codebook is subsequently employed to quantize the query image for producing compact, discriminative, and scalable descriptors. As q-Codebook may be simultaneously generated at both sides, without transmitting the entire vocabulary, only small overhead (e.g. codebook ID and codeword 0/1 index) is incurred to reconstruct the query signature at the server end. To fulfill m-Codebooks and q-Codebook, we adopt a Bi-layer Sparse Coding method to learn the sparse relationships of codewords vs. codebooks as well as codebooks vs. query images via l1 regularization. Experiments on benchmarking datasets have demonstrated the extremely small descriptor's supervior performance in image retrieval.
Jie Lin 0001, Ling-Yu Duan, Jie Chen 0006, Rongrong Ji, Siwei Luo, Wen Gao 0001
ICASSP1
2012 Learning sparse tag patterns for social image classification
abstract
User-generated tags associated with images from social media (e.g., Flickr) provide valuable textual resources for image classification. However, the noisy and huge tag vocabulary heavily degrades the effectiveness and efficiency of state-of-the-art image classification methods that exploited auxiliary web data. To alleviate the problem, we introduce a Sparse Tag Patterns (STP) model to discover sparsity constrained co-occurrence tag patterns from large scale user contributed tags among social data. To fulfill the compactness and discriminability, we formulate STP as a problem of minimizing a quadratic loss function regularized by the bi-layer l1norm. We treat the learned STP as alternative intermediate semantic image feature and verify its superiority within a search-based image classification framework. Experiments on 240K social images associated with millions of tags have demonstrated encouraging performance of the proposed method compared to the state-of-the-art.
Jie Lin 0001, Ling-Yu Duan, Junsong Yuan 0001, Qingyong Li, Siwei Luo
ICIP1
2012 Social Image Tagging by Mining Sparse Tag Patterns from Auxiliary Data
abstract
User-given tags associated with social images from photosharing websites (e.g., Flickr) are valuable auxiliary resources for the image tagging task. However, social images often suffer from noisy and incomplete tags, heavily degrading the effectiveness of previous image tagging approaches. To alleviate the problem, we introduce a Sparse Tag Patterns (STP) model to discover noiseless and complementary cooccurrence tag patterns from large scale user contributed tags among auxiliary web data. To fulfill the compactness and discriminability, we formulate the STP model as a problem of minimizing quadratic loss function regularized by bi-layer ℓ1norm. We treat the learned STP as a universal knowledge base and verify its superiority within a data-driven image tagging framework. Experimental results over 1 million auxiliary data demonstrate superior performance of the proposed method compared to the state-of-the-art.
Jie Lin 0001, Junsong Yuan 0001, Ling-Yu Duan, Siwei Luo, Wen Gao 0001
ICME1