Jianxin Wu 0001

dblp:w/JianxinWu · DBLP profile ↗
← Back
135ranked-venue papers
22as first author
43since 2021 · last 2026
0000-0002-2085-7568ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 100 · 19 first-author · 39 since 2021Graphics, computer vision, multimedia, augmented reality and games · 73 · 9 first-author · 25 since 2021Computer networks · 5Databases, data management, data science and information retrieval · 5 · 1 first-author · 2 since 2021Systems, architecture and hardware · 4 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 DTL: Parameter- and Memory-Efficient Disentangled Vision Learning
abstract
The cost of finetuning a pretrained model on downstream tasks steadily increases as they grow larger. Parameter-efficient transfer learning (PETL) is proposed to reduce this cost by changing only a tiny subset of trainable parameters. But, the GPU memory footprint during training is not effectively reduced in PETL. This issue happens because trainable parameters from these methods are generally tightly entangled with the backbone, such that a lot of intermediate states have to be stored for back propagation. To alleviate this issue, we introduce Disentangled Transfer Learning (DTL), which disentangles the trainable parameters from the backbone using a lightweight Compact Side Network (CSN). By progressively extracting task-specific information with a few low-rank linear mappings and appropriately adding the information back to the backbone, CSN effectively realizes knowledge transfer in various downstream recognition tasks. We further extend DTL to more difficult tasks such as object detection and semantic segmentation by employing a more sparse architectural design. Extensive experiments validate the effectiveness of DTL, which not only reduces a large amount of GPU memory usage and trainable parameters, but also outperforms existing PETL methods by a significant margin in accuracy.
Minghao Fu 0001, Zonghao Ding, Jianxin Wu 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2025 Treasures in Discarded Weights for LLM Quantization
abstract
In recent years, large language models (LLMs) have developed rapidly and revolutionized natural language processing. However, high storage overhead and computing costs limit LLM deployment in resource-constrained environments. Quantization algorithms can effectively compress LLMs and accelerate inference, but they lead to loss in precision, especially in low-bit scenarios. In this paper, we find that the discarded weight values caused by quantization in fact contain treasures to improve LLMs' accuracy. To excavate those hidden treasures, we construct search spaces around these discarded weights and those weights within the search space can seamlessly be incorporated into the original quantization weights. To determine which weights should be merged, we design a plug-and-play weight compensation framework to capture global information and keep the weights with the highest potential benefits. Our framework can be combined with various LLM quantization algorithms to achieve higher precision without additional inference overhead. We validate the effectiveness of our approach on widely used benchmark datasets for LLMs.
Hao Yu 0027, Bohua Chen, Zelan Yang, Jianxin Wu 0001
AAAI7
2025 All You Need in Knowledge Distillation Is a Tailored Coordinate System
abstract
Knowledge Distillation (KD) is essential in transferring dark knowledge from a large teacher to a small student network, such that the student can be much more efficient than the teacher but with comparable accuracy. Existing KD methods, however, rely on a large teacher trained specifically for the target task, which is both very inflexible and inefficient. In this paper, we argue that a SSL-pretrained model can effectively act as the teacher and its dark knowledge can be captured by the coordinate system or linear subspace where the features lie in. We then need only one forward pass of the teacher, and then tailor the coordinate system (TCS) for the student network. Our TCS method is teacher-free and applies to diverse architectures, works well for KD and practical few-shot learning, allows cross-architecture distillation with large capacity gap. Experiments show that TCS achieves significantly higher accuracy than state-of-the-art KD methods, while only requiring roughly half of their training time and GPU memory costs.
Jianxin Wu 0001
AAAI3
2025 Quantization without Tears
abstract
Deep neural networks, while achieving remarkable success across diverse tasks, demand significant resources, including computation, GPU memory, bandwidth, storage, and energy. Network quantization, as a standard compression and acceleration technique, reduces storage costs and enables potential inference acceleration by discretizing network weights and activations into a finite set of integer values. However, current quantization methods are often complex and sensitive, requiring extensive task-specific hyperparameters, where even a single misconfiguration can impair model performance, limiting generality across different models and tasks. In this paper, we propose Quantization without Tears (QwT), a method that simultaneously achieves quantization speed, accuracy, simplicity, and generality. The key insight of QwT is to incorporate a lightweight additional structure into the quantized network to mitigate information loss during quantization. This structure consists solely of a small set of linear layers, keeping the method simple and efficient. More importantly, it provides a closed-form solution, allowing us to improve accuracy effortlessly under 2 minutes. Extensive experiments across various vision, language, and multimodal tasks demonstrate that QwT is both highly effective and versatile. In fact, our approach offers a robust solution for network quantization that combines simplicity, accuracy, and adaptability, which provides new insights for the design of novel quantization paradigms.
Minghao Fu 0001, Hao Yu 0027, Jie Shao 0001, Jianxin Wu 0001
CVPR6
2025 Minimal Interaction Seperated Tuning: A New Paradigm for Visual Adaptation
abstract
The rapid scaling of large vision pretrained models makes fine-tuning tasks more and more difficult on devices with low computational resources. We explore a new visual adaptation paradigm called separated tuning, which treats large pretrained models as standalone feature extractors that run on powerful cloud servers. The fine-tuning carries out on devices which possess only low computational resources (slow CPU, no GPU, small memory, etc.) Existing methods that are potentially suitable for our separated tuning paradigm are discussed. But, three major drawbacks hinder their application in separated tuning: low adaptation capability, large adapter network, and in particular, high information transfer overhead. To address these issues, we propose Minimal Interaction Separated Tuning, or MIST, which reveals that the sum of intermediate features from pretrained models not only has minimal information transfer but also has high adaptation capability. With a lightweight attention-based adaptor network, MIST achieves information transfer efficiency, parameter efficiency, computational and memory efficiency, and at the same time demonstrates competitive results on various visual adaptation benchmarks.
Ningyuan Tang, Minghao Fu 0001, Jianxin Wu 0001
CVPR3
2025 Memory-Efficient Generative Models via Product Quantization
Jie Shao 0001, Hao Yu 0027, Jianxin Wu 0001
ICCV4
2025 GPLQ: A General, Practical, and Lightning QAT Method for Vision Transformers
abstract
Vision Transformers (ViTs) are essential in computer vision but are computationally intensive, too. Model quantization, particularly to low bit-widths like 4-bit, aims to alleviate this difficulty, yet existing Post-Training Quantization (PTQ) and Quantization-Aware Training (QAT) methods exhibit significant limitations. PTQ often incurs substantial accuracy drop, while QAT achieves high accuracy but suffers from prohibitive computational costs, limited generalization to downstream tasks, training instability, and lacking of open-source codebase. To address these challenges, this paper introduces General, Practical, and Lightning Quantization (GPLQ), a novel framework designed for efficient and effective ViT quantization. GPLQ is founded on two key empirical insights: the paramount importance of activation quantization and the necessity of preserving the model's original optimization basin to maintain generalization. Consequently, GPLQ employs a sequential activation-first, weights-later strategy. Stage 1 keeps weights in FP32 while quantizing activations with a feature mimicking loss in only 1 epoch to keep it stay in the same basin, thereby preserving generalization. Stage 2 quantizes weights using a PTQ method. As a result, GPLQ is 100x faster than existing QAT methods, lowers memory footprint to levels even below FP32 training, and achieves 4-bit model performance that is highly competitive with FP32 models in terms of both accuracy on ImageNet and generalization to diverse downstream tasks, including fine-grained visual classification and object detection. We release an easy-to-use open-source toolkit supporting multiple vision tasks at [GPLQ code](https://github.com/wujx2001/GPLQ).
Guang Liang, Xinyao Liu, Jianxin Wu 0001
NeurIPS3
2025 Who Reasons in the Large Language Models?
abstract
Despite the impressive performance of large language models (LLMs), the process of endowing them with new capabilities---such as mathematical reasoning---remains largely empirical and opaque. A critical open question is whether reasoning abilities stem from the entire model, specific modules, or are merely artifacts of overfitting. In this work, we hypothesize that the reasoning capabilities in well-trained LLMs are primarily attributed to the output projection module (o_proj) in the Transformer’s multi-head self-attention (MHSA) module. To support this hypothesis, we introduce Stethoscope for Networks (SfN), a suite of diagnostic tools designed to probe and analyze the internal behaviors of LLMs. Using SfN, we provide both circumstantial and empirical evidence suggesting that o_proj plays a central role in enabling reasoning, whereas other modules contribute more to fluent dialogue. These findings offer a new perspective on LLM interpretability and open avenues for more targeted training strategies, potentially enabling more efficient and specialized LLMs.
Jie Shao 0001, Jianxin Wu 0001
NeurIPS2
2025 Dual DTL: Double-Side Compact Transfer Networks for Efficient ViT Adaptation
Zonghao Ding, Minghao Fu 0001, Jianxin Wu 0001
PRCV (2)3
2025 Reviving undersampling for long-tailed learning
Hao Yu 0027, Yingxiao Du, Jianxin Wu 0001
Pattern Recognit.3
2025 Coarse is better? A new pipeline towards self-supervised learning with uncurated images
Yin-Yin He, Jianxin Wu 0001
Pattern Recognit.3
2024 DTL: Disentangled Transfer Learning for Visual Recognition
abstract
When pre-trained models become rapidly larger, the cost of fine-tuning on downstream tasks steadily increases, too. To economically fine-tune these models, parameter-efficient transfer learning (PETL) is proposed, which only tunes a tiny subset of trainable parameters to efficiently learn quality representations. However, current PETL methods are facing the dilemma that during training the GPU memory footprint is not effectively reduced as trainable parameters. PETL will likely fail, too, if the full fine-tuning encounters the out-of-GPU-memory issue. This phenomenon happens because trainable parameters from these methods are generally entangled with the backbone, such that a lot of intermediate states have to be stored in GPU memory for gradient propagation. To alleviate this problem, we introduce Disentangled Transfer Learning (DTL), which disentangles the trainable parameters from the backbone using a lightweight Compact Side Network (CSN). By progressively extracting task-specific information with a few low-rank linear mappings and appropriately adding the information back to the backbone, CSN effectively realizes knowledge transfer in various downstream tasks. We conducted extensive experiments to validate the effectiveness of our method. The proposed method not only reduces a large amount of GPU memory usage and trainable parameters, but also outperforms existing PETL methods by a significant margin in accuracy, achieving new state-of-the-art on several standard benchmarks.
Minghao Fu 0001, Jianxin Wu 0001
AAAI3
2024 Rectify the Regression Bias in Long-Tailed Object Detection
Minghao Fu 0001, Jie Shao 0001, Jianxin Wu 0001
ECCV (28)5
2024 Unified Low-rank Compression Framework for Click-through Rate Prediction
abstract
Deep Click-Through Rate (CTR) prediction models play an important role in modern industrial recommendation scenarios. However, high memory overhead and computational costs limit their deployment in resource-constrained environments. Low-rank approximation is an effective method for computer vision and natural language processing models, but its application in compressing CTR prediction models has been less explored. Due to the limited memory and computing resources, compression of CTR prediction models often confronts three fundamental challenges, i.e., (1). How to reduce the model sizes to adapt to edge devices? (2). How to speed up CTR prediction model inference? (3). How to retain the capabilities of original models after compression? Previous low-rank compression research mostly uses tensor decomposition, which can achieve a high parameter compression ratio, but brings in AUC degradation and additional computing overhead. To address these challenges, we propose a unified low-rank decomposition framework for compressing CTR prediction models. We find that even with the most classic matrix decomposition SVD method, our framework can achieve better performance than the original model. To further improve the effectiveness of our framework, we locally compress the output features instead of compressing the model weights. Our unified low-rank compression framework can be applied to embedding tables and MLP layers in various CTR prediction models. Extensive experiments on two academic datasets and one real industrial benchmark demonstrate that, with 3--5× model size reduction, our compressed models can achieve both faster inference and higher AUC than the uncompressed original models. Our code is at https://github.com/yuhao318/Atomic_Feature_Mimicking.
Hao Yu 0027, Minghao Fu 0001, Jiandong Ding, Jianxin Wu 0001
KDD5
2024 DiffuLT: Diffusion for Long-tail Recognition Without External Knowledge
abstract
This paper introduces a novel pipeline for long-tail (LT) recognition that diverges from conventional strategies. Instead, it leverages the long-tailed dataset itself to generate a balanced proxy dataset without utilizing external data or model. We deploy a diffusion model trained from scratch on only the long-tailed dataset to create this proxy and verify the effectiveness of the data produced. Our analysis identifies approximately-in-distribution (AID) samples, which slightly deviate from the real data distribution and incorporate a blend of class information, as the crucial samples for enhancing the generative model's performance in long-tail classification. We promote the generation of AID samples during the training of a generative model by utilizing a feature extractor to guide the process and filter out detrimental samples during generation. Our approach, termed Diffusion model for Long-Tail recognition (DiffuLT), represents a pioneer application of generative models in long-tail recognition. DiffuLT achieves state-of-the-art results on CIFAR10-LT, CIFAR100-LT, and ImageNet-LT, surpassing leading competitors by significant margins. Comprehensive ablations enhance the interpretability of our pipeline. Notably, the entire generative process is conducted without relying on external data or pre-trained model weights, which leads to its generalizability to real-world long-tailed scenarios.
Jie Shao 0001, Jianxin Wu 0001
NeurIPS4
2024 Tobias: A Random CNN Sees Objects
abstract
This paper starts by revealing a surprising finding: without any learning, a randomly initialized CNN can localize objects surprisingly well. That is, a CNN has an inductive bias to naturally focus on objects, named as Tobias (“Theobjectisatsight”) in this paper. This empirical inductive bias is further theoretically analyzed and empirically verified, and successfully applied to self-supervised learning as well as supervised learning. For self-supervised learning, a CNN is encouraged to learn representations that focus on the foreground object, by transforming every image into various versions with different backgrounds, where the foreground and background separation is guided by Tobias. Experimental results show that the proposed Tobias significantly improves downstream tasks, especially for object detection. This paper also shows that Tobias has consistent improvements on training sets of different sizes, and is more resilient to changes in image augmentations. Furthermore, we apply Tobias to supervised image classification by letting the average pooling layer focus on foreground regions, which achieves improved performance on various benchmarks.
Yun-Hao Cao, Jianxin Wu 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2024 Practical Network Acceleration With Tiny Sets: Hypothesis, Theory, and Algorithm
abstract
Due to data privacy issues, accelerating networks with tiny training sets has become a critical need in practice. Previous methods achieved promising results empirically by filter-level pruning. In this paper, we both study this problem theoretically and propose an effective algorithm aligning well with our theoretical results. First, we propose the finetune convexity hypothesis to explain why recent few-shot compression algorithms do not suffer from overfitting problems. Based on it, a theory is further established to explain these methods for the first time. Compared to naively finetuning a pruned network, feature mimicking is proved to achieve a lower variance of parameters and hence enjoys easier optimization. With our theoretical conclusions, we claim dropping blocks is a fundamentally superior few-shot compression scheme in terms of more convex optimization and a higher acceleration ratio. To choose which blocks to drop, we propose a new metric, recoverability, to effectively measure the difficulty of recovering the compressed network. Finally, we propose an algorithm namedPractiseto accelerate networks using only tiny sets of training images.Practiseoutperforms previous methods by a significant margin. For 22% latency reduction,Practisesurpasses previous methods by on average 7 percentage points on ImageNet-1k. It also enjoys high generalization ability, working well under data-free or out-of-domain data settings, too.
Guo-Hua Wang, Jianxin Wu 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 Compressing Transformers: Features Are Low-Rank, but Weights Are Not!
abstract
Transformer and its variants achieve excellent results in various computer vision and natural language processing tasks, but high computational costs and reliance on large training datasets restrict their deployment in resource-constrained settings. Low-rank approximation of model weights has been effective in compressing CNN models, but its application to transformers has been less explored and is less effective. Existing methods require the complete dataset to fine-tune compressed models, which are both time-consuming and data-hungry. This paper reveals that the features (i.e., activations) are low-rank, but model weights are surprisingly not low-rank. Hence, AAFM is proposed, which adaptively determines the compressed model structure and locally compresses each linear layer's output features rather than the model weights. A second stage, GFM, optimizes the entire compressed network holistically. Both AAFM and GFM only use few training samples without labels, that is, they are few-shot, unsupervised, fast and effective. For example, with only 2K images without labels, 33% of the parameters are removed in DeiT-B with 18.8% relative throughput increase, but only a 0.23% accuracy loss for ImageNet recognition. The proposed methods are successfully applied to the language modeling task in NLP, too. Besides, the few-shot compressed models generalize well in downstream tasks.
Hao Yu 0027, Jianxin Wu 0001
AAAI2
2023 Quantized Feature Distillation for Network Quantization
abstract
Neural network quantization aims to accelerate and trim full-precision neural network models by using low bit approximations. Methods adopting the quantization aware training (QAT) paradigm have recently seen a rapid growth, but are often conceptually complicated. This paper proposes a novel and highly effective QAT method, quantized feature distillation (QFD). QFD first trains a quantized (or binarized) representation as the teacher, then quantize the network using knowledge distillation (KD). Quantitative results show that QFD is more flexible and effective (i.e., quantization friendly) than previous quantization methods. QFD surpasses existing methods by a noticeable margin on not only image classification but also object detection, albeit being much simpler. Furthermore, QFD quantizes ViT and Swin-Transformer on MS-COCO detection and segmentation, which verifies its potential in real world deployment. To the best of our knowledge, this is the first time that vision transformers have been quantized in object detection and image segmentation tasks.
Yin-Yin He, Jianxin Wu 0001
AAAI3
2023 No One Left Behind: Improving the Worst Categories in Long-Tailed Learning
abstract
Unlike the case when using a balanced training dataset, the per-class recall (i.e., accuracy) of neural networks trained with an imbalanced dataset are known to vary a lot from category to category. The convention in long-tailed recognition is to manually split all categories into three subsets and report the average accuracy within each subset. We argue that under such an evaluation setting, some categories are inevitably sacrificed. On one hand, focusing on the average accuracy on a balanced test set incurs little penalty even if some worst performing categories have zero accuracy. On the other hand, classes in the “Few” subset do not necessarily perform worse than those in the “Many” or “Medium” subsets. We therefore advocate to focus more on improving the lowest recall among all categories and the harmonic mean of all recall values. Specifically, we propose a simple plug-in method that is applicable to a wide range of methods. By simply retraining the classifier of an existing pretrained model with our proposed loss function and using an optional ensemble trick that combines the predictions of the two classifiers, we achieve a more uniform distribution of recall values across categories, which leads to a higher harmonic mean accuracy while the (arithmetic) average accuracy is still high. The effectiveness of our method is justified on widely used benchmark datasets.
Yingxiao Du, Jianxin Wu 0001
CVPR2
2023 Practical Network Acceleration with Tiny Sets
abstract
Due to data privacy issues, accelerating networks with tiny training sets has become a critical need in practice. Previous methods mainly adopt filter-level pruning to accelerate networks with scarce training samples. In this paper, we reveal that dropping blocks is a fundamentally superior approach in this scenario. It enjoys a higher acceleration ratio and results in a better latency-accuracy performance under the few-shot setting. To choose which blocks to drop, we propose a new concept namely recoverability to measure the difficulty of recovering the compressed network. Our recoverability is efficient and effective for choosing which blocks to drop. Finally, we propose an algorithm named Practise to accelerate networks using only tiny sets of training images. Practise outperforms previous methods by a significant margin. For 22% latency reduction, Practise surpasses previous methods by on average 7% on ImageNet-1k. It also enjoys high generalization ability, working well under data-free or out-of-domain data settings, too. Our code is at https://github.com/DoctorKey/Practise.
Guo-Hua Wang, Jianxin Wu 0001
CVPR2
2023 Multi-Label Self-Supervised Learning with Scene Images
abstract
Self-supervised learning (SSL) methods targeting scene images have seen a rapid growth recently, and they mostly rely on either a dedicated dense matching mechanism or a costly unsupervised object discovery module. This paper shows that instead of hinging on these strenuous operations, quality image representations can be learned by treating scene/multi-label image SSL simply as a multi-label classification problem, which greatly simplifies the learning framework. Specifically, multiple binary pseudo-labels are assigned for each input image by comparing its embeddings with those in two dictionaries, and the network is optimized using the binary cross entropy loss. The proposed method is named Multi-Label Self-supervised learning (MLS). Visualizations qualitatively show that clearly the pseudo-labels by MLS can automatically find semantically similar pseudo-positive pairs across different images to facilitate contrastive learning. MLS learns high quality representations on MS-COCO and achieves state-of-the-art results on classification, detection and segmentation benchmarks. At the same time, MLS is much simpler than existing methods, making it easier to deploy and for further exploration.
Minghao Fu 0001, Jianxin Wu 0001
ICCV3
2023 A unified pruning framework for vision transformers
Hao Yu 0027, Jianxin Wu 0001
Sci. China Inf. Sci.2
2023 Salvage of Supervision in Weakly Supervised Object Detection and Segmentation
abstract
Weakly supervised vision tasks, including detection and segmentation, have attracted much attention in the vision community recently. However, the lack of detailed and precise annotations in the weakly supervised case leads to a large accuracy gap between weakly- and fully-supervised methods. In this article, we propose a new framework, Salvage of Supervision (SoS), with the key idea being to effectively harness every potentially useful supervisory signal in weakly supervised vision tasks. Starting with weakly supervised object detection (WSOD), we propose SoS-WSOD to shrink the technology gap between WSOD and FSOD, which utilizes the weak image-level labels, the pseudo-labels, and the power of semi-supervised object detection for WSOD. Moreover, SoS-WSOD removes restrictions in traditional WSOD methods, including the reliance on ImageNet pretraining and inability to use modern backbones. The SoS framework also extends to weakly supervised semantic segmentation and instance segmentation. On several weakly supervised vision benchmarks, SoS achieves significant performance boost and generalization ability.
Lin Sui, Chen-Lin Zhang, Jianxin Wu 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2023 Weakly supervised foreground learning for weakly supervised localization and detection
Chen-Lin Zhang, Yin Li 0003, Jianxin Wu 0001
Pattern Recognit.3
2022 A Random CNN Sees Objects: One Inductive Bias of CNN and Its Applications
abstract
This paper starts by revealing a surprising finding: without any learning, a randomly initialized CNN can localize objects surprisingly well. That is, a CNN has an inductive bias to naturally focus on objects, named as Tobias ("The object is at sight") in this paper. This empirical inductive bias is further analyzed and successfully applied to self-supervised learning (SSL). A CNN is encouraged to learn representations that focus on the foreground object, by transforming every image into various versions with different backgrounds, where the foreground and background separation is guided by Tobias. Experimental results show that the proposed Tobias significantly improves downstream tasks, especially for object detection. This paper also shows that Tobias has consistent improvements on training sets of different sizes, and is more resilient to changes in image augmentations.
Yun-Hao Cao, Jianxin Wu 0001
AAAI2
2022 Salvage of Supervision in Weakly Supervised Object Detection
abstract
Weakly supervised object detection (WSOD) has recently attracted much attention. However, the lack of bounding-box supervision makes its accuracy much lower than fully supervised object detection (FSOD), and currently modern FSOD techniques cannot be applied to WSOD. To bridge the performance and technical gaps between WSOD and FSOD, this paper proposes a new framework, Salvage of Supervision (SoS), with the key idea being to harness every potentially useful supervisory signal in WSOD: the weak image-level labels, the pseudo-labels, and the power of semi-supervised object detection. This paper proposes new approaches to utilize these weak and noisy signals effectively, and shows that each type of supervisory signal brings in notable improvements, outperforms existing WSOD methods (which mainly use only the weak labels) by large margins. The proposed SoS- WSOD method also has the ability to freely use modern FSOD techniques. SoS-WSOD achieves 64.4 mAP50on VOC2007, 61.9 mAP50on VOC2012 and 16.6 mAP50:95on MS-COCO, and also has fast inference speed. Ablations and visualization further verify the effectiveness of SoS.
Lin Sui, Chen-Lin Zhang, Jianxin Wu 0001
CVPR3
2022 Compressing Models with Few Samples: Mimicking then Replacing
abstract
Few-sample compression aims to compress a big redundant model into a small compact one with only few samples. If we fine-tune models with these limited few samples directly, models will be vulnerable to overfit and learn almost nothing. Hence, previous methods optimize the compressed model layer-by-layer and try to make every layer have the same outputs as the corresponding layer in the teacher model, which is cumbersome. In this paper, we propose a new framework named Mimicking then Replacing (MiR) for few-sample compression, which firstly urges the pruned model to output the same features as the teacher's in the penultimate layer, and then replaces teacher's layers before penultimate with a well-tuned compact one. Unlike previous layer-wise reconstruction methods, our MiR optimizes the entire network holistically, which is not only simple and effective, but also unsupervised and general. MiR outperforms previous methods with large margins. Codes is available at https://github.com/cjnjuwhy/MiR.
Junjie Liu 0003, Xin Ma 0031, Yang Yong, Zhenhua Chai, Jianxin Wu 0001
CVPR6
2022 Synergistic Self-supervised and Quantization Learning
Yun-Hao Cao, Peiqin Sun, Yechang Huang, Jianxin Wu 0001, Shuchang Zhou 0001
ECCV (30)4
2022 Training Vision Transformers with only 2040 Images
Yun-Hao Cao, Hao Yu 0027, Jianxin Wu 0001
ECCV (25)3
2022 Worst Case Matters for Few-Shot Recognition
Minghao Fu 0001, Yun-Hao Cao, Jianxin Wu 0001
ECCV (20)3
2022 ActionFormer: Localizing Moments of Actions with Transformers
Chen-Lin Zhang, Jianxin Wu 0001, Yin Li 0003
ECCV (4)2
2022 Distilling Knowledge by Mimicking Features
abstract
Knowledge distillation (KD) is a popular method to train efficient networks ("student") with the help of high-capacity networks ("teacher"). Traditional methods use the teacher's soft logits as extra supervision to train the student network. In this paper, we argue that it is more advantageous to make the student mimic the teacher's features in the penultimate layer. Not only the student can directly learn more effective information from the teacher feature, feature mimicking can also be applied for teachers trained without a softmax layer. Experiments show that it can achieve higher accuracy than traditional KD. To further facilitate feature mimicking, we decompose a feature vector into the magnitude and the direction. We argue that the teacher should give more freedom to the student feature's magnitude, and let the student pay more attention on mimicking the feature direction. To meet this requirement, we propose a loss term based on locality-sensitive hashing (LSH). With the help of this new loss, our method indeed mimics feature directions more accurately, relaxes constraints on feature magnitudes, and achieves state-of-the-art distillation accuracy. We provide theoretical analyses of how LSH facilitates feature direction mimicking, and further extend feature mimicking to multi-label recognition and object detection.
Guo-Hua Wang, Yifan Ge, Jianxin Wu 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2022 Fine-Grained Image Analysis With Deep Learning: A Survey
abstract
Fine-grained image analysis (FGIA) is a longstanding and fundamental problem in computer vision and pattern recognition, and underpins a diverse set of real-world applications. The task of FGIA targets analyzing visual objects from subordinate categories, e.g., species of birds or models of cars. The small inter-class and large intra-class variation inherent to fine-grained image analysis makes it a challenging problem. Capitalizing on advances in deep learning, in recent years we have witnessed remarkable progress in deep learning powered FGIA. In this paper we present a systematic survey of these advances, where we attempt to re-define and broaden the field of FGIA by consolidating two fundamental fine-grained research areas - fine-grained image recognition and fine-grained image retrieval. In addition, we also review other key issues of FGIA, such as publicly available benchmark datasets and related domain-specific applications. We conclude by highlighting several research directions and open problems which need further exploration from the community.
Xiu-Shen Wei, Yi-Zhe Song, Oisin Mac Aodha, Jianxin Wu 0001, Yuxin Peng 0001, Jinhui Tang 0001, Jian Yang 0003, Serge J. Belongie
IEEE Trans. Pattern Anal. Mach. Intell.4
2022 Versatile, full-spectrum, and swift network sampling for model generation
Jianxin Wu 0001
Pattern Recognit.3
2022 ECML: An Ensemble Cascade Metric-Learning Mechanism Toward Face Verification
abstract
Face verification can be regarded as a two-class fine-grained visual-recognition problem. Enhancing the feature's discriminative power is one of the key problems to improve its performance. Metric-learning technology is often applied to address this need while achieving a good tradeoff between underfitting, and overfitting plays a vital role in metric learning. Hence, we propose a novel ensemble cascade metric-learning (ECML) mechanism. In particular, hierarchical metric learning is executed in a cascade way to alleviate underfitting. Meanwhile, at each learning level, the features are split into nonoverlapping groups. Then, metric learning is executed among the feature groups in the ensemble manner to resist overfitting. Considering the feature distribution characteristics of faces, a robust Mahalanobis metric-learning method (RMML) with a closed-form solution is additionally proposed. It can avoid the computation failure issue on an inverse matrix faced by some well-known metric-learning approaches (e.g., KISSME). Embedding RMML into the proposed ECML mechanism, our metric-learning paradigm (EC-RMML) can run in the one-pass learning manner. The experimental results demonstrate that EC-RMML is superior to state-of-the-art metric-learning methods for face verification. The proposed ECML mechanism is also applicable to other metric-learning approaches.
Fu Xiong, Yang Xiao 0007, Zhiguo Cao 0001, Yancheng Wang 0002, Joey Tianyi Zhou, Jianxin Wu 0001
IEEE Trans. Cybern.6
2021 Bag of Tricks for Long-Tailed Visual Recognition with Deep Convolutional Neural Networks
abstract
In recent years, visual recognition on challenging long-tailed distributions, where classes often exhibit extremely imbalanced frequencies, has made great progress mostly based on various complex paradigms (e.g., meta learning). Apart from these complex methods, simple refinements on training procedures also make contributions. These refinements, also called tricks, are minor but effective, such as adjustments in the data distribution or loss functions. However, different tricks might conflict with each other. If users apply these long-tail related tricks inappropriately, it could cause worse recognition accuracy than expected. Unfortunately, there has not been a scientific guideline of these tricks in the literature. In this paper, we first collect existing tricks in long-tailed visual recognition and then perform extensive and systematic experiments, in order to give a detailed experimental guideline and obtain an effective combination of these tricks. Furthermore, we also propose a novel data augmentation approach based on class activation maps for long-tailed recognition, which can be friendly combined with re-sampling methods and shows excellent results. By assembling these tricks scientifically, we can outperform state-of-the-art methods on four long-tailed benchmark datasets, including ImageNet-LT and iNaturalist 2018. Our code is open-source and available at https://github.com/zhangyongshun/BagofTricks-LT.
Xiu-Shen Wei, Boyan Zhou, Jianxin Wu 0001
AAAI4
2021 Distilling Virtual Examples for Long-tailed Recognition
abstract
We tackle the long-tailed visual recognition problem from the knowledge distillation perspective by proposing a Distill the Virtual Examples (DiVE) method. Specifically, by treating the predictions of a teacher model as virtual examples, we prove that distilling from these virtual examples is equivalent to label distribution learning under certain constraints. We show that when the virtual example distribution becomes flatter than the original input distribution, the under-represented tail classes will receive significant improvements, which is crucial in long-tailed recognition. The proposed DiVE method can explicitly tune the virtual example distribution to become flat. Extensive experiments on three benchmark datasets, including the large-scale iNaturalist ones, justify that the proposed DiVE method can significantly outperform state-of-the-art methods. Further-more, additional analyses and experiments verify the virtual example interpretation, and demonstrate the effectiveness of tailored designs in DiVE for long-tailed problems.
Yin-Yin He, Jianxin Wu 0001, Xiu-Shen Wei
ICCV2
2021 Webly Supervised Fine-Grained Recognition: Benchmark Datasets and An Approach
abstract
Learning from the web can ease the extreme dependence of deep learning on large-scale manually labeled datasets. Especially for fine-grained recognition, which targets at distinguishing subordinate categories, it will significantly reduce the labeling costs by leveraging free web data. Despite its significant practical and research value, the webly supervised fine-grained recognition problem is not extensively studied in the computer vision community, largely due to the lack of high-quality datasets. To fill this gap, in this paper we construct two new benchmark webly supervised fine-grained datasets, termed WebFG-496 and WebiNat-5089, respectively. In concretely, WebFG-496 consists of three sub-datasets containing a total of 53,339 web training images with 200 species of birds (Web-bird), 100 types of aircrafts (Web-aircraft), and 196 models of cars (Web-car). For WebiNat-5089, it contains 5089 sub-categories and more than 1.1 million web training images, which is the largest webly supervised fine-grained dataset ever. As a minor contribution, we also propose a novel webly supervised method (termed "Peer-learning") for benchmarking these datasets. Comprehensive experimental results and analyses on two new benchmark datasets demonstrate that the proposed method achieves superior performance over the competing baseline models and states-of-the-art. Our benchmark datasets and the source codes of Peer-learning have been made available at https://github.com/NUST-Machine-Intelligence-Laboratory/weblyFG-dataset.
Zeren Sun, Yazhou Yao, Xiu-Shen Wei, Fumin Shen, Jianxin Wu 0001, Jian Zhang 0002, Heng Tao Shen
ICCV6
2021 Residual Attention: A Simple but Effective Method for Multi-Label Recognition
abstract
Multi-label image recognition is a challenging computer vision task of practical use. Progresses in this area, how-ever, are often characterized by complicated methods, heavy computations, and lack of intuitive explanations. To effectively capture different spatial regions occupied by objects from different categories, we propose an embarrassingly simple module, named class-specific residual attention (CSRA). CSRA generates class-specific features for every category by proposing a simple spatial attention score, and then combines it with the class-agnostic average pooling feature. CSRA achieves state-of-the-art results on multi-label recognition, and at the same time is much simpler than them. Furthermore, with only 4 lines of code, CSRA also leads to consistent improvement across many diverse pretrained models and datasets without any extra training. CSRA is both easy to implement and light in computations, which also enjoys intuitive explanations and visualizations.
Jianxin Wu 0001
ICCV2
2021 Mixup Without Hesitation
Hao Yu 0027, Jianxin Wu 0001
ICIG (2)3
2021 Neural random subspace
Yun-Hao Cao, Jianxin Wu 0001, Hanchen Wang 0002, Joan Lasenby
Pattern Recognit.2
2021 Multi-Instance Learning With Emerging Novel Class
abstract
Diverse applications involving complicated data objects such as proteins and images are solved by applying multi-instance learning (MIL) algorithms. However, few MIL algorithms can deal with problems in an open and dynamic environment, where new categories of samples emerge. In this type of emerging novel class setting, algorithms should be able to not only classify the samples from the observed classes accurately, but also recognize the samples from the novel class. In this paper, we focus on the Multi-Instance learning with Emerging Novel class (MIEN) problem, and formulate MIEN from a metric learning perspective. We extract key instances to form the “super-bag” for each observed class, and non-key instances from all the observed classes to form a “meta super-bag”. Based on these super-bags, we propose the MIEN-metric method to learn discriminative metrics for classifying MIL bags from the observed classes and recognizing bags from the novel class. Experimental results of diverse domains, e.g., biological function annotation, text categorization, and object-centric/scene-centric image classification, show MIEN-metric outperforms other baseline methods significantly when the novel class emerges. Meanwhile, MIEN-metric is comparable with state-of-the-art MIL algorithms for binary classification in the traditional MIL setting.
Xiu-Shen Wei, Han-Jia Ye, Xin Mu, Jianxin Wu 0001, Chunhua Shen, Zhi-Hua Zhou
IEEE Trans. Knowl. Data Eng.4
2020 Repetitive Reprediction Deep Decipher for Semi-Supervised Learning
abstract
Most recent semi-supervised deep learning (deep SSL) methods used a similar paradigm: use network predictions to update pseudo-labels and use pseudo-labels to update network parameters iteratively. However, they lack theoretical support and cannot explain why predictions are good candidates for pseudo-labels. In this paper, we propose a principled end-to-end framework named deep decipher (D2) for SSL. Within the D2 framework, we prove that pseudo-labels are related to network predictions by an exponential link function, which gives a theoretical support for using predictions as pseudo-labels. Furthermore, we demonstrate that updating pseudo-labels by network predictions will make them uncertain. To mitigate this problem, we propose a training strategy called repetitive reprediction (R2). Finally, the proposed R2-D2 method is tested on the large-scale ImageNet dataset and outperforms state-of-the-art methods by 5 percentage points.
Guo-Hua Wang, Jianxin Wu 0001
AAAI2
2020 Deep Discriminative CNN with Temporal Ensembling for Ambiguously-Labeled Image Classification
abstract
In this paper, we study the problem of image classification where training images are ambiguously annotated with multiple candidate labels, among which only one is correct but is not accessible during the training phase. Due to the adopted non-deep framework and improper disambiguation strategies, traditional approaches are usually short of the representation ability and discrimination ability, so their performances are still to be improved. To remedy these two shortcomings, this paper proposes a novel approach termed “Deep Discriminative CNN” (D2CNN) with temporal ensembling. Specifically, to improve the representation ability, we innovatively employ the deep convolutional neural networks for ambiguously-labeled image classification, in which the well-known ResNet is adopted as our backbone. To enhance the discrimination ability, we design an entropy-based regularizer to maximize the margin between the potentially correct label and the unlikely ones of each image. In addition, we utilize the temporally assembled predictions of different epochs to guide the training process so that the latent groundtruth label can be confidently highlighted. This is much superior to the traditional disambiguation operations which treat all candidate labels equally and identify the hidden groundtruth label via some heuristic ways. Thorough experimental results on multiple datasets firmly demonstrate the effectiveness of our proposed D2CNN when compared with other existing state-of-the-art approaches.
Jiehui Deng, Xiuhua Chen, Chen Gong 0002, Jianxin Wu 0001, Jian Yang 0003
AAAI5
2020 Neural Network Pruning With Residual-Connections and Limited-Data
abstract
Filter level pruning is an effective method to accelerate the inference speed of deep CNN models. Although numerous pruning algorithms have been proposed, there are still two open issues. The first problem is how to prune residual connections. We propose to prune both channels inside and outside the residual connections via a KL-divergence based criterion. The second issue is pruning with limited data. We observe an interesting phenomenon: directly pruning on a small dataset is usually worse than fine-tuning a small model which is pruned or trained from scratch on the large dataset. Knowledge distillation is an effective approach to compensate for the weakness of limited data. However, the logits of a teacher model may be noisy. In order to avoid the influence of label noise, we propose a label refinement approach to solve this problem. Experiments have demonstrated the effectiveness of our method (CURL, Compression Using Residual-connections and Limited-data). CURL significantly outperforms previous state-of-the-art methods on ImageNet. More importantly, when pruning on small datasets, CURL achieves comparable or much better performance than fine-tuning a pretrained small model.
Jian-Hao Luo, Jianxin Wu 0001
CVPR2
2020 Rethinking the Route Towards Weakly Supervised Object Localization
abstract
Weakly supervised object localization (WSOL) aims to localize objects with only image-level labels. Previous methods often try to utilize feature maps and classification weights to localize objects using image level annotations indirectly. In this paper, we demonstrate that weakly supervised object localization should be divided into two parts: class-agnostic object localization and object classification. For class-agnostic object localization, we should use class-agnostic methods to generate noisy pseudo annotations and then perform bounding box regression on them without class labels. We propose the pseudo supervised object localization (PSOL) method as a new way to solve WSOL. Our PSOL models have good transferability across different datasets without fine-tuning. With generated pseudo bounding boxes, we achieve 58.00% localization accuracy on ImageNet and 74.74% localization accuracy on CUB-200, which have a large edge over previous models.
Chen-Lin Zhang, Yun-Hao Cao, Jianxin Wu 0001
CVPR3
2020 AutoPruner: An end-to-end trainable filter pruning method for efficient deep model inference
Jian-Hao Luo, Jianxin Wu 0001
Pattern Recognit.2
2020 Simultaneous 3D hand detection and pose estimation using single depth images
Yu Zhang 0004, Siya Mi, Jianxin Wu 0001, Xin Geng 0001
Pattern Recognit. Lett.3
2019 Probabilistic End-To-End Noise Correction for Learning With Noisy Labels
abstract
Deep learning has achieved excellent performance in various computer vision tasks, but requires a lot of training examples with clean labels. It is easy to collect a dataset with noisy labels, but such noise makes networks overfit seriously and accuracies drop dramatically. To address this problem, we propose an end-to-end framework called PENCIL, which can update both network parameters and label estimations as label distributions. PENCIL is independent of the backbone network structure and does not need an auxiliary clean dataset or prior information about noise, thus it is more general and robust than existing methods and is easy to apply. PENCIL outperformed previous state-of-the-art methods by large margins on both synthetic and real-world datasets with different noise types and noise rates. Experiments show that PENCIL is robust on clean datasets, too.
Jianxin Wu 0001
CVPR2
2019 CodeAttention: translating source code to comments by exploiting the code constructs
Wenhao Zheng 0001, Ming Li 0005, Jianxin Wu 0001
Frontiers Comput. Sci.4
2019 ThiNet: Pruning CNN Filters for a Thinner Net
abstract
This paper aims at accelerating and compressing deep neural networks to deploy CNN models into small devices like mobile phones or embedded gadgets. We focus on filter level pruning, i.e., the whole filter will be discarded if it is less important. An effective and unified framework, ThiNet (stands for "Thin Net"), is proposed in this paper. We formally establish filter pruning as an optimization problem, and reveal that we need to prune filters based on statistics computed from its next layer, not the current layer, which differentiates ThiNet from existing methods. We also propose "gcos" (Group COnvolution with Shuffling), a more accurate group convolution scheme, to further reduce the pruned model size. Experimental results demonstrate the effectiveness of our method, which has advanced the state-of-the-art. Moreover, we show that the original VGG-16 model can be compressed into a very small model (ThiNet-Tiny) with only 2.66 MB model size, but still preserve AlexNet level accuracy. This small model is evaluated on several benchmarks with different vision tasks (e.g., classification, detection, segmentation), and shows excellent generalization ability.
Jian-Hao Luo, Hao Zhang 0038, Chen-Wei Xie, Jianxin Wu 0001, Weiyao Lin
IEEE Trans. Pattern Anal. Mach. Intell.5
2019 Unsupervised object discovery and co-localization by deep descriptor transformation
Xiu-Shen Wei, Chen-Lin Zhang, Jianxin Wu 0001, Chunhua Shen, Zhi-Hua Zhou
Pattern Recognit.3
2019 Improving CNN linear layers with power mean non-linearity
Chen-Lin Zhang, Jianxin Wu 0001
Pattern Recognit.2
2019 Piecewise Classifier Mappings: Learning Fine-Grained Learners for Novel Categories With Few Examples
abstract
Humans are capable of learning a new fine-grained concept with very little supervision, e.g., few exemplary images for a species of bird, yet our best deep learning systems need hundreds or thousands of labeled examples. In this paper, we try to reduce this gap by studying the fine-grained image recognition problem in a challenging few-shot learning setting, termed few-shot fine-grained recognition (FSFG). The task of FSFG requires the learning systems to build classifiers for the novel fine-grained categories from few examples (only one or less than five). To solve this problem, we propose an end-to-end trainable deep network, which is inspired by the state-of-the-art fine-grained recognition model and is tailored for the FSFG task. Specifically, our network consists of a bilinear feature learning module and a classifier mapping module: while the former encodes the discriminative information of an exemplar image into a feature vector, the latter maps the intermediate feature into the decision boundary of the novel category. The key novelty of our model is a "piecewise mappings" function in the classifier mapping module, which generates the decision boundary via learning a set of more attainable sub-classifiers in a more parameter-economic way. We learn the exemplar-to-classifier mapping based on an auxiliary dataset in a meta-learning fashion, which is expected to be able to generalize to novel categories. By conducting comprehensive experiments on three fine-grained datasets, we demonstrate that the proposed method achieves superior performance over the competing baselines.
Xiu-Shen Wei, Peng Wang 0023, Lingqiao Liu, Chunhua Shen, Jianxin Wu 0001
IEEE Trans. Image Process.5
2018 Action Recognition With Coarse-to-Fine Deep Feature Integration and Asynchronous Fusion
abstract
Action recognition is an important yet challenging task in computer vision. In this paper, we propose a novel deep-based framework for action recognition, which improves the recognition accuracy by: 1) deriving more precise features for representing actions, and 2) reducing the asynchrony between different information streams. We first introduce a coarse-to-fine network which extracts shared deep features at different action class granularities and progressively integrates them to obtain a more accurate feature representation for input actions. We further introduce an asynchronous fusion network. It fuses information from different streams by asynchronously integrating stream-wise features at different time points, hence better leveraging the complementary information in different streams. Experimental results on action recognition benchmarks demonstrate that our approach achieves the state-of-the-art performance.
Weiyao Lin, Ke Lu 0002, Bin Sheng 0001, Jianxin Wu 0001, Bingbing Ni, Hongkai Xiong
AAAI5
2018 Coarse-to-Fine: A RNN-Based Hierarchical Attention Model for Vehicle Re-identification
Xiu-Shen Wei, Chen-Lin Zhang, Lingqiao Liu, Chunhua Shen, Jianxin Wu 0001
ACCV (2)5
2018 Age Estimation Using Expectation of Label Distribution Learning
abstract
Age estimation performance has been greatly improved by using convolutional neural network. However, existing methods have an inconsistency between the training objectives and evaluation metric, so they may be suboptimal. In addition, these methods always adopt image classification or face recognition models with a large amount of parameters, which bring expensive computation cost and storage overhead. To alleviate these issues, we design a lightweight network architecture and propose a unified framework which can jointly learn age distribution and regress age. The effectiveness of our approach has been demonstrated on apparent and real age estimation tasks. Our method achieves new state-of-the-art results using the single model with 36$\times$ fewer parameters and 2.6$\times$ reduction in inference time. Moreover, our method can achieve comparable results as the state-of-the-art even though model parameters are further reduced to 0.9M~(3.8MB disk storage). We also analyze that Ranking methods are implicitly learning label distributions.
Bin-Bin Gao, Jianxin Wu 0001, Xin Geng 0001
IJCAI3
2018 Mask-CNN: Localizing parts and selecting descriptors for fine-grained bird species categorization
Xiu-Shen Wei, Chen-Wei Xie, Jianxin Wu 0001, Chunhua Shen
Pattern Recognit.3
2018 Deep Bimodal Regression of Apparent Personality Traits from Short Video Sequences
abstract
Apparent personality analysis (APA) is an important problem of personality computing, and furthermore, automatic APA becomes a hot and challenging topic in computer vision and multimedia. In this paper, we propose a deep learning solution to APA from short video sequences. In order to capture rich information from both the visual and audio modality of videos, we tackle these tasks with our Deep Bimodal Regression (DBR) framework. In DBR, for the visual modality, we modify the traditional convolutional neural networks for exploiting important visual cues. In addition, taking into account the model efficiency, we extract audio representations and build a linear regressor for the audio modality. For combining the complementary information from the two modalities, we ensemble these predicted regression scores by both early fusion and late fusion. Finally, based on the proposed framework, we come up with a solution for the Apparent Personality Analysis competition track in the ChaLearn Looking at People challenge in association with ECCV 2016. Our DBR is the winner (first place) of this challenge with 86 registered participants. Beyond the competition, we further investigate the performance of different loss functions in our visual models, and prove non-convex loss functions for regression are optimal on the human-labeled video data.
Xiu-Shen Wei, Chen-Lin Zhang, Hao Zhang 0038, Jianxin Wu 0001
IEEE Trans. Affect. Comput.4
2017 Sunrise or Sunset: Selective Comparison Learning for Subtle Attribute Recognition
Bin-Bin Gao, Jianxin Wu 0001
BMVC3
2017 ThiNet: A Filter Level Pruning Method for Deep Neural Network Compression
abstract
We propose an efficient and unified framework, namely ThiNet, to simultaneously accelerate and compress CNN models in both training and inference stages. We focus on the filter level pruning, i.e., the whole filter would be discarded if it is less important. Our method does not change the original network structure, thus it can be perfectly supported by any off-the-shelf deep learning libraries. We formally establish filter pruning as an optimization problem, and reveal that we need to prune filters based on statistics information computed from its next layer, not the current layer, which differentiates ThiNet from existing methods. Experimental results demonstrate the effectiveness of this strategy, which has advanced the state-of-the-art. We also show the performance of ThiNet on ILSVRC-12 benchmark. ThiNet achieves 3.31 x FLOPs reduction and 16.63× compression on VGG-16, with only 0.52% top-5 accuracy drop. Similar experiments with ResNet-50 reveal that even for a compact network, ThiNet can also reduce more than half of the parameters and FLOPs, at the cost of roughly 1% top-5 accuracy drop. Moreover, the original VGG-16 model can be further pruned into a very small model with only 5.05MB model size, preserving AlexNet level accuracy but showing much stronger generalization ability.
Jian-Hao Luo, Jianxin Wu 0001, Weiyao Lin
ICCV2
2017 Adaptive Feeding: Achieving Fast and Accurate Detections by Adaptively Combining Object Detectors
abstract
Object detection aims at high speed and accuracy simultaneously. However, fast models are usually less accurate, while accurate models cannot satisfy our need for speed. A fast model can be 10 times faster but 50% less accurate than an accurate model. In this paper, we propose Adaptive Feeding (AF) to combine a fast (but less accurate) detector and an accurate (but slow) detector, by adaptively determining whether an image is easy or hard and choosing an appropriate detector for it. In practice, we build a cascade of detectors, including the AF classifier which make the easy vs. hard decision and the two detectors. The AF classifier can be tuned to obtain different tradeoff between speed and accuracy, which has negligible training time and requires no additional training data. Experimental results on the PASCAL VOC, MS COCO and Caltech Pedestrian datasets confirm that AF has the ability to achieve comparable speed as the fast detector and comparable accuracy as the accurate one at the same time. As an example, by combining the fast SSD300 with the accurate SSD500 detector, AF leads to 50% speedup over SSD500 with the same precision on the VOC2007 test set.
Bin-Bin Gao, Jianxin Wu 0001
ICCV3
2017 Deep Descriptor Transforming for Image Co-Localization
abstract
Reusable model design becomes desirable with the rapid expansion of machine learning applications. In this paper, we focus on the reusability of pre-trained deep convolutional models. Specifically, different from treating pre-trained models as feature extractors, we reveal more treasures beneath convolutional layers, i.e., the convolutional activations could act as a detector for the common object in the image co-localization problem. We propose a simple but effective method, named Deep Descriptor Transforming (DDT), for evaluating the correlations of descriptors and then obtaining the category-consistent regions, which can accurately locate the common object in a set of images. Empirical studies validate the effectiveness of the proposed DDT method. On benchmark image co-localization datasets, DDT consistently outperforms existing state-of-the-art methods by a large margin. Moreover, DDT also demonstrates good generalization ability for unseen categories and robustness for dealing with noisy data.
Xiu-Shen Wei, Chen-Lin Zhang, Yao Li 0003, Chen-Wei Xie, Jianxin Wu 0001, Chunhua Shen, Zhi-Hua Zhou
IJCAI5
2017 Image categorization with resource constraints: introduction, challenges and advances
Jian-Hao Luo, Jianxin Wu 0001
Frontiers Comput. Sci.3
2017 Structured Learning of Binary Codes with Column Generation for Optimizing Ranking Measures
Guosheng Lin, Fayao Liu, Chunhua Shen, Jianxin Wu 0001, Heng Tao Shen
Int. J. Comput. Vis.4
2017 Sparse multiple instance learning as document classification
Shengye Yan, Jianxin Wu 0001
Multim. Tools Appl.4
2017 A Tube-and-Droplet-Based Approach for Representing and Analyzing Motion Trajectories
abstract
Trajectory analysis is essential in many applications. In this paper, we address the problem of representing motion trajectories in a highly informative way, and consequently utilize it for analyzing trajectories. Our approach first leverages the complete information from given trajectories to construct a thermal transfer field which provides a context-rich way to describe the global motion pattern in a scene. Then, a 3D tube is derived which depicts an input trajectory by integrating its surrounding motion patterns contained in the thermal transfer field. The 3D tube effectively: 1) maintains the movement information of a trajectory, 2) embeds the complete contextual motion pattern around a trajectory, 3) visualizes information about a trajectory in a clear and unified way. We further introduce a droplet-based process. It derives a droplet vector from a 3D tube, so as to characterize the high-dimensional 3D tube information in a simple but effective way. Finally, we apply our tube-and-droplet representation to trajectory analysis applications including trajectory clustering, trajectory classification & abnormality detection, and 3D action recognition. Experimental comparisons with state-of-the-art algorithms demonstrate the effectiveness of our approach.
Weiyao Lin, Hongteng Xu, Junchi Yan, Mingliang Xu 0001, Jianxin Wu 0001, Zicheng Liu 0001
IEEE Trans. Pattern Anal. Mach. Intell.6
2017 Editorial of the Special Issue on Multi-instance Learning in Pattern Recognition and Vision
Jianxin Wu 0001, Xiang Bai, Marco Loog, Fabio Roli, Zhi-Hua Zhou
Pattern Recognit.1
2017 Deep Label Distribution Learning With Label Ambiguity
abstract
Convolutional neural networks (ConvNets) have achieved excellent recognition performance in various visual recognition tasks. A large labeled training set is one of the most important factors for its success. However, it is difficult to collect sufficient training images with precise labels in some domains, such as apparent age estimation, head pose estimation, multilabel classification, and semantic segmentation. Fortunately, there is ambiguous information among labels, which makes these tasks different from traditional classification. Based on this observation, we convert the label of each image into a discrete label distribution, and learn the label distribution by minimizing a Kullback-Leibler divergence between the predicted and ground-truth label distributions using deep ConvNets. The proposed deep label distribution learning (DLDL) method effectively utilizes the label ambiguity in both feature learning and classifier learning, which help prevent the network from overfitting even when the training set is small. Experimental results show that the proposed approach produces significantly better results than the state-of-the-art methods for age estimation and head pose estimation. At the same time, it also improves recognition performance for multi-label classification and semantic segmentation tasks.
Bin-Bin Gao, Chen-Wei Xie, Jianxin Wu 0001, Xin Geng 0001
IEEE Trans. Image Process.4
2017 Learning Correspondence Structures for Person Re-Identification
abstract
This paper addresses the problem of handling spatial misalignments due to camera-view changes or human-pose variations in person re-identification. We first introduce a boosting-based approach to learn a correspondence structure, which indicates the patchwise matching probabilities between images from a target camera pair. The learned correspondence structure can not only capture the spatial correspondence pattern between cameras but also handle the viewpoint or human-pose variation in individual images. We further introduce a global constraint-based matching process. It integrates a global matching constraint over the learned correspondence structure to exclude cross-view misalignments during the image patch matching process, hence achieving a more reliable matching score between images. Finally, we also extend our approach by introducing a multi-structure scheme, which learns a set of local correspondence structures to capture the spatial correspondence sub-patterns between a camera pair, so as to handle the spatial misalignments between individual images in a more precise way. Experimental results on various data sets demonstrate the effectiveness of our approach.
Weiyao Lin, Junchi Yan, Mingliang Xu 0001, Jianxin Wu 0001, Jingdong Wang 0001, Ke Lu 0002
IEEE Trans. Image Process.5
2017 Selective Convolutional Descriptor Aggregation for Fine-Grained Image Retrieval
abstract
Deep convolutional neural network models pre-trained for the ImageNet classification task have been successfully adopted to tasks in other domains, such as texture description and object proposal generation, but these tasks require annotations for images in the new domain. In this paper, we focus on a novel and challenging task in the pure unsupervised setting: fine-grained image retrieval. Even with image labels, fine-grained images are difficult to classify, letting alone the unsupervised retrieval task. We propose the selective convolutional descriptor aggregation (SCDA) method. The SCDA first localizes the main object in fine-grained images, a step that discards the noisy background and keeps useful deep descriptors. The selected descriptors are then aggregated and the dimensionality is reduced into a short feature vector using the best practices we found. The SCDA is unsupervised, using no image label or bounding box annotation. Experiments on six fine-grained data sets confirm the effectiveness of the SCDA for fine-grained image retrieval. Besides, visualization of the SCDA features shows that they correspond to visual attributes (even subtle ones), which might explain SCDA's high-mean average precision in fine-grained retrieval. Moreover, on general image retrieval data sets, the SCDA achieves comparable retrieval results with the state-of-the-art general image retrieval approaches.
Xiu-Shen Wei, Jian-Hao Luo, Jianxin Wu 0001, Zhi-Hua Zhou
IEEE Trans. Image Process.3
2017 Scalable Algorithms for Multi-Instance Learning
abstract
Multi-instance learning (MIL) has been widely applied to diverse applications involving complicated data objects, such as images and genes. However, most existing MIL algorithms can only handle small- or moderate-sized data. In order to deal with large-scale MIL problems, we propose MIL based on the vector of locally aggregated descriptors representation (miVLAD) and MIL based on the Fisher vector representation (miFV), two efficient and scalable MIL algorithms. They map the original MIL bags into new vector representations using their corresponding mapping functions. The new feature representations keep essential bag-level information, and at the same time lead to excellent MIL performances even when linear classifiers are used. Thanks to the low computational cost in the mapping step and the scalability of linear classifiers, miVLAD and miFV can handle large-scale MIL data efficiently and effectively. Experiments show that miVLAD and miFV not only achieve comparable accuracy rates with the state-of-the-art MIL algorithms, but also have hundreds of times faster speed. Moreover, we can regard the new miVLAD and miFV representations as multiview data, which improves the accuracy rates in most cases. In addition, our algorithms perform well even when they are used without parameter tuning (i.e., adopting the default parameters), which is convenient for practical MIL applications.
Xiu-Shen Wei, Jianxin Wu 0001, Zhi-Hua Zhou
IEEE Trans. Neural Networks Learn. Syst.2
2016 Representing Sets of Instances for Visual Recognition
abstract
In computer vision, a complex entity such as an image or video is often represented as a set of instance vectors, which are extracted from different parts of that entity. Thus, it is essential to design a representation to encode information in a set of instances robustly. Existing methods such as FV and VLAD are designed based on a generative perspective, and their performances fluctuate when difference types of instance vectors are used (i.e., they are not robust). The proposed D3 method effectively compares two sets as two distributions, and proposes a directional total variation distance (DTVD) to measure their dissimilarity. Furthermore, a robust classifier-based method is proposed to estimate DTVD robustly, and to efficiently represent these sets. D3 is evaluated in action and image recognition tasks. It achieves excellent robustness, accuracy and speed.
Jianxin Wu 0001, Bin-Bin Gao
AAAI1
2016 Exploit Bounding Box Annotations for Multi-Label Object Recognition
abstract
Convolutional neural networks (CNNs) have shown great performance as general feature representations for object recognition applications. However, for multi-label images that contain multiple objects from different categories, scales and locations, global CNN features are not optimal. In this paper, we incorporate local information to enhance the feature discriminative power. In particular, we first extract object proposals from each image. With each image treated as a bag and object proposals extracted from it treated as instances, we transform the multi-label recognition problem into a multi-class multi-instance learning problem. Then, in addition to extracting the typical CNN feature representation from each proposal, we propose to make use of ground-truth bounding box annotations (strong labels) to add another level of local information by using nearest-neighbor relationships of local regions to form a multi-view pipeline. The proposed multi-view multiinstance framework utilizes both weak and strong labels effectively, and more importantly it has the generalization ability to even boost the performance of unseen categories by partial strong labels from other categories. Our framework is extensively compared with state-of-the-art handcrafted feature based methods and CNN based methods on two multi-label benchmark datasets. The experimental results validate the discriminative power and the generalization ability of the proposed framework. With strong labels, our framework is able to achieve state-of-the-art results in both datasets.
Hao Yang 0033, Joey Tianyi Zhou, Yu Zhang 0004, Bin-Bin Gao, Jianxin Wu 0001, Jianfei Cai 0001
CVPR5
2016 Learning compact binary codes from higher-order tensors via Free-Form Reshaping and Binarized Multilinear PCA
abstract
For big, high-dimensional dense features, it is important to learn compact binary codes or compress them for greater memory efficiency. This paper proposes a Binarized Multilinear PCA (BMP) method for this problem with Free-Form Reshaping (FFR) of such features to higher-order tensors, lifting the structure-modelling restriction in traditional tensor models. The reshaped tensors are transformed to a subspace using multilinear PCA. Then, we unsupervisedly select features and supervisedly binarize them with a minimum-classification-error scheme to get compact binary codes. We evaluate BMP on two scene recognition datasets against state-of-the-art algorithms. The FFR works well in experiments. With the same number of compression parameters (model size), BMP has much higher classification accuracy. To achieve the same accuracy or compression ratio, BMP has an order of magnitude smaller number of compression parameters. Thus, BMP has great potential in memory-sensitive applications such as mobile computing and big data analytics.
Haiping Lu, Jianxin Wu 0001, Yu Zhang 0004
IJCNN2
2016 Good Practices for Learning to Recognize Actions Using FV and VLAD
abstract
High dimensional representations such as Fisher vectors (FV) and vectors of locally aggregated descriptors (VLAD) have shown state-of-the-art accuracy for action recognition in videos. The high dimensionality, on the other hand, also causes computational difficulties when scaling up to large-scale video data. This paper makes three lines of contributions to learning to recognize actions using high dimensional representations. First, we reviewed several existing techniques that improve upon FV or VLAD in image classification, and performed extensive empirical evaluations to assess their applicability for action recognition. Our analyses of these empirical results show that normality and bimodality are essential to achieve high accuracy. Second, we proposed a new pooling strategy for VLAD and three simple, efficient, and effective transformations for both FV and VLAD. Both proposed methods have shown higher accuracy than the original FV/VLAD method in extensive evaluations. Third, we proposed and evaluated new feature selection and compression methods for the FV and VLAD representations. This strategy uses only 4% of the storage of the original representation, but achieves comparable or even higher accuracy. Based on these contributions, we recommend a set of good practices for action recognition in videos for practitioners in this field.
Jianxin Wu 0001, Yu Zhang 0004, Weiyao Lin
IEEE Trans. Cybern.1
2016 Action Recognition in Still Images With Minimum Annotation Efforts
abstract
We focus on the problem of still image-based human action recognition, which essentially involves making prediction by analyzing human poses and their interaction with objects in the scene. Besides image-level action labels (e.g., riding, phoning), during both training and testing stages, existing works usually require additional input of human bounding boxes to facilitate the characterization of the underlying human-object interactions. We argue that this additional input requirement might severely discourage potential applications and is not very necessary. To this end, a systematic approach was developed in this paper to address this challenging problem of minimum annotation efforts, i.e., to perform recognition in the presence of only image-level action labels in the training stage. Experimental results on three benchmark data sets demonstrate that compared with the state-of-the-art methods that have privileged access to additional human bounding-box annotations, our approach achieves comparable or even superior recognition accuracy using only action annotations in training. Interestingly, as a by-product in many cases, our approach is able to segment out the precise regions of underlying human-object interactions.
Yu Zhang 0004, Li Cheng 0001, Jianxin Wu 0001, Jianfei Cai 0001, Minh N. Do, Jiangbo Lu
IEEE Trans. Image Process.3
2016 Weakly Supervised Fine-Grained Categorization With Part-Based Image Representation
abstract
In this paper, we propose a fine-grained image categorization system with easy deployment. We do not use any object/part annotation (weakly supervised) in the training or in the testing stage, but only class labels for training images. Fine-grained image categorization aims to classify objects with only subtle distinctions (e.g., two breeds of dogs that look alike). Most existing works heavily rely on object/part detectors to build the correspondence between object parts, which require accurate object or object part annotations at least for training images. The need for expensive object annotations prevents the wide usage of these methods. Instead, we propose to generate multi-scale part proposals from object proposals, select useful part proposals, and use them to compute a global image representation for categorization. This is specially designed for the weakly supervised fine-grained categorization task, because useful parts have been shown to play a critical role in existing annotation-dependent works, but accurate part detectors are hard to acquire. With the proposed image representation, we can further detect and visualize the key (most discriminative) parts in objects of different classes. In the experiments, the proposed weakly supervised method achieves comparable or better accuracy than the state-of-the-art weakly supervised methods and most existing annotation-dependent methods on three challenging datasets. Its success suggests that it is not always necessary to learn expensive object/part detectors in fine-grained image categorization.
Yu Zhang 0004, Xiu-Shen Wei, Jianxin Wu 0001, Jianfei Cai 0001, Jiangbo Lu, Minh N. Do
IEEE Trans. Image Process.3
2016 A Diffusion and Clustering-Based Approach for Finding Coherent Motions and Understanding Crowd Scenes
abstract
This paper addresses the problem of detecting coherent motions in crowd scenes and presents its two applications in crowd scene understanding: semantic region detection and recurrent activity mining. It processes input motion fields (e.g., optical flow fields) and produces a coherent motion field named thermal energy field. The thermal energy field is able to capture both motion correlation among particles and the motion trends of individual particles, which are helpful to discover coherency among them. We further introduce a two-step clustering process to construct stable semantic regions from the extracted time-varying coherent motions. These semantic regions can be used to recognize pre-defined activities in crowd scenes. Finally, we introduce a cluster-and-merge process, which automatically discovers recurrent activities in crowd scenes by clustering and merging the extracted coherent motions. Experiments on various videos demonstrate the effectiveness of our approach.
Weiyao Lin, Yang Mi, Weiyue Wang 0002, Jianxin Wu 0001, Jingdong Wang 0001, Tao Mei 0001
IEEE Trans. Image Process.4
2016 Compact Representation of High-Dimensional Feature Vectors for Large-Scale Image Recognition and Retrieval
abstract
In large-scale visual recognition and image retrieval tasks, feature vectors, such as Fisher vector (FV) or the vector of locally aggregated descriptors (VLAD), have achieved state-of-the-art results. However, the combination of the large numbers of examples and high-dimensional vectors necessitates dimensionality reduction, in order to reduce its storage and CPU costs to a reasonable range. In spite of the popularity of various feature compression methods, this paper shows that the feature (dimension) selection is a better choice for high-dimensional FV/VLAD than the feature (dimension) compression methods, e.g., product quantization. We show that strong correlation among the feature dimensions in the FV and the VLAD may not exist, which renders feature selection a natural choice. We also show that, many dimensions in FV/VLAD are noise. Throwing them away using feature selection is better than compressing them and useful dimensions altogether using feature compression methods. To choose features, we propose an efficient importance sorting algorithm considering both the supervised and unsupervised cases, for visual recognition and image retrieval, respectively. Combining with the 1-bit quantization, feature selection has achieved both higher accuracy and less computational cost than feature compression methods, such as product quantization, on the FV and the VLAD image representations.
Yu Zhang 0004, Jianxin Wu 0001, Jianfei Cai 0001
IEEE Trans. Image Process.2
2015 Person Re-Identification with Correspondence Structure Learning
abstract
This paper addresses the problem of handling spatial misalignments due to camera-view changes or human-pose variations in person re-identification. We first introduce a boosting-based approach to learn a correspondence structure which indicates the patch-wise matching probabilities between images from a target camera pair. The learned correspondence structure can not only capture the spatial correspondence pattern between cameras but also handle the viewpoint or human-pose variation in individual images. We further introduce a global-based matching process. It integrates a global matching constraint over the learned correspondence structure to exclude cross-view misalignments during the image patch matching process, hence achieving a more reliable matching score between images. Experimental results on various datasets demonstrate the effectiveness of our approach.
Weiyao Lin, Junchi Yan, Jianxin Wu 0001, Jingdong Wang 0001
ICCV5
2015 GROPING: Geomagnetism and cROwdsensing Powered Indoor NaviGation
abstract
Although a large number of WiFi fingerprinting based indoor localization systems have been proposed, our field experience with Google Maps Indoor (GMI), the only system available for public testing, shows that it is far from mature for indoor navigation. In this paper, we first report our field studies with GMI, as well as experiment results aiming to explain our unsatisfactory GMI experience. Then motivated by the obtained insights, we propose GROPING as a self-contained indoor navigation system independent of any infrastructural support. GROPING relies on geomagnetic fingerprints that are far more stable than WiFi fingerprints, and it exploits crowdsensing to construct floor maps rather than expecting individual venues to supply digitized maps. Based on our experiments with 20 participants in various floors of a big shopping mall, GROPING is able to deliver a sufficient accuracy for localization and thus provides smooth navigation experience.
Chi Zhang 0064, Kalyan Subbu, Jun Luo 0001, Jianxin Wu 0001
IEEE Trans. Mob. Comput.4
2015 Linear Regression-Based Efficient SVM Learning for Large-Scale Classification
abstract
For large-scale classification tasks, especially in the classification of images, additive kernels have shown a state-of-the-art accuracy. However, even with the recent development of fast algorithms, learning speed and the ability to handle large-scale tasks are still open problems. This paper proposes algorithms for large-scale support vector machines (SVM) classification and other tasks using additive kernels. First, a linear regression SVM framework for general nonlinear kernel is proposed using linear regression to approximate gradient computations in the learning process. Second, we propose a power mean SVM (PmSVM) algorithm for all additive kernels using nonsymmetric explanatory variable functions. This nonsymmetric kernel approximation has advantages over the existing methods: 1) it does not require closed-form Fourier transforms and 2) it does not require extra training for the approximation either. Compared on benchmark large-scale classification data sets with millions of examples or millions of dense feature dimensions, PmSVM has achieved the highest learning speed and highest accuracy among recent algorithms in most cases.
Jianxin Wu 0001, Hao Yang 0033
IEEE Trans. Neural Networks Learn. Syst.1
2015 Automatic Face Naming by Learning Discriminative Affinity Matrices From Weakly Labeled Images
abstract
Given a collection of images, where each image contains several faces and is associated with a few names in the corresponding caption, the goal of face naming is to infer the correct name for each face. In this paper, we propose two new methods to effectively solve this problem by learning two discriminative affinity matrices from these weakly labeled images. We first propose a new method called regularized low-rank representation by effectively utilizing weakly supervised information to learn a low-rank reconstruction coefficient matrix while exploring multiple subspace structures of the data. Specifically, by introducing a specially designed regularizer to the low-rank representation method, we penalize the corresponding reconstruction coefficients related to the situations where a face is reconstructed by using face images from other subjects or by using itself. With the inferred reconstruction coefficient matrix, a discriminative affinity matrix can be obtained. Moreover, we also develop a new distance metric learning method called ambiguously supervised structural metric learning by using weakly supervised information to seek a discriminative distance metric. Hence, another discriminative affinity matrix can be obtained using the similarity matrix (i.e., the kernel matrix) based on the Mahalanobis distances of the data. Observing that these two affinity matrices contain complementary information, we further combine them to obtain a fused affinity matrix, based on which we develop a new iterative scheme to infer the name of each face. Comprehensive experiments demonstrate the effectiveness of our approach.
Shijie Xiao, Dong Xu 0001, Jianxin Wu 0001
IEEE Trans. Neural Networks Learn. Syst.3
2014 Learning with Augmented Multi-Instance View
Yue Zhu 0001, Jianxin Wu 0001, Yuan Jiang 0001, Zhi-Hua Zhou
ACML2
2014 Towards Good Practices for Action Video Encoding
abstract
High dimensional representations such as VLAD or FV have shown excellent accuracy in action recognition. This paper shows that a proper encoding built upon VLAD can achieve further accuracy boost with only negligible computational cost. We empirically evaluated various VLAD improvement technologies to determine good practices in VLAD-based video encoding. Furthermore, we propose an interpretation that VLAD is a maximum entropy linear feature learning process. Combining this new perspective with observed VLAD data distribution properties, we propose a simple, lightweight, but powerful bimodal encoding method. Evaluated on 3 benchmark action recognition datasets (UCF101, HMDB51 and Youtube), the bimodal encoding improves VLAD by large margins in action recognition.
Jianxin Wu 0001, Yu Zhang 0004, Weiyao Lin
CVPR1
2014 Compact Representation for Image Classification: To Choose or to Compress?
abstract
In large scale image classification, features such as Fisher vector or VLAD have achieved state-of-the-art results. However, the combination of large number of examples and high dimensional vectors necessitates dimensionality reduction, in order to reduce its storage and CPU costs to a reasonable range. In spite of the popularity of various feature compression methods, this paper argues that feature selection is a better choice than feature compression. We show that strong multicollinearity among feature dimensions may not exist, which undermines feature compression's effectiveness and renders feature selection a natural choice. We also show that many dimensions are noise and throwing them away is helpful for classification. We propose a supervised mutual information (MI) based importance sorting algorithm to choose features. Combining with 1-bit quantization, MI feature selection has achieved both higher accuracy and less computational cost than feature compression methods such as product quantization and BPBC.
Yu Zhang 0004, Jianxin Wu 0001, Jianfei Cai 0001
CVPR2
2014 A Dual-Sensor Enabled Indoor Localization System with Crowdsensing Spot Survey
abstract
We present MaWi - a smart phone based scalable indoor localization system. Central to MaWi is a novel framework combining two self-contained but complementary localization techniques: Wi-Fi and Ambient Magnetic Field. Combining the two techniques, MaWi not only achieves a high localization accuracy, but also effectively reduces human labor in building fingerprint databases: to avoid war-driving, MaWi is designed to work with low quality fingerprint databases that can be efficiently built by only one person. Our experiments demonstrate that MaWi, with a fingerprint database as scarce as one data sample at each spot, outperforms the state-of-the-art proposals working on a richer fingerprint database.
Chi Zhang 0064, Jun Luo 0001, Jianxin Wu 0001
DCOSS3
2014 Optimizing Ranking Measures for Compact Binary Code Learning
Guosheng Lin, Chunhua Shen, Jianxin Wu 0001
ECCV (3)3
2014 Finding Coherent Motions and Semantic Regions in Crowd Scenes: A Diffusion and Clustering Approach
Weiyue Wang 0002, Weiyao Lin, Yuanzhe Chen, Jianxin Wu 0001, Jingdong Wang 0001, Bin Sheng 0001
ECCV (1)4
2014 Scalable Multi-instance Learning
abstract
Multi-instance learning (MIL) has been widely applied to diverse applications involving complicated data objects such as images and genes. However, most existing MIL algorithms can only handle small-or moderate-sized data. In order to deal with the large scale problems in MIL, we propose an efficient and scalable MIL algorithm named miFV. Our algorithm maps the original MIL bags into a new feature vector representation, which can obtain bag-level information, and meanwhile lead to excellent performances even with linear classifiers. In consequence, thanks to the low computational cost in the mapping step and the scalability of linear classifiers, miFV can handle large scale MIL data efficiently and effectively. Experiments show that miFV not only achieves comparable accuracy rates with state-of-the-art MIL algorithms, but has hundreds of times faster speed than other MIL algorithms.
Xiu-Shen Wei, Jianxin Wu 0001, Zhi-Hua Zhou
ICDM2
2014 Poster abstract: MaWi: a hybrid magnetic and wi-fi system for scalable indoor localization
Chi Zhang 0064, Jun Luo 0001, Jianxin Wu 0001
IPSN3
2014 Representing And Recognizing Motion Trajectories: A Tube And Droplet Approach
abstract
This paper addresses the problem of representing and recognizing motion trajectories. We first propose to derive scene-related equipotential lines for points in a motion trajectory and concatenate them to construct a 3D tube for representing the trajectory. Based on this 3D tube, a droplet-based method is further proposed which derives a "water droplet" from the 3D tube and recognizes trajectory activities accordingly. Our proposed 3D tube can effectively embed both motion and scene-related information of a motion trajectory while the proposed droplet- based method can suitably catch the characteristics of the 3D tube for activity recognition. Experimental results demonstrate the effectiveness of our approach.
Weiyao Lin, Hang Su 0006, Jianxin Wu 0001, Jinjun Wang, Yu Zhou 0015
ACM Multimedia4
2014 Facial expression cloning with elastic and muscle models
Weiyao Lin, Bing Zhou 0003, Zhenzhong Chen 0001, Bin Sheng 0001, Jianxin Wu 0001
J. Vis. Commun. Image Represent.6
2014 A New Network-Based Algorithm for Human Activity Recognition in Videos
abstract
In this paper, a new network-transmission-based (NTB) algorithm is proposed for human activity recognition in videos. The proposed NTB algorithm models the entire scene as an error-free network. In this network, each node corresponds to a patch of the scene and each edge represents the activity correlation between the corresponding patches. Based on this network, we further model people in the scene as packages, while human activities can be modeled as the process of package transmission in the network. By analyzing these specific package transmission processes, various activities can be effectively detected. The implementation of our NTB algorithm into abnormal activity detection and group activity recognition are described in detail in this paper. Experimental results demonstrate the effectiveness of our proposed algorithm.
Weiyao Lin, Yuanzhe Chen, Jianxin Wu 0001, Hanli Wang, Bin Sheng 0001, Hongxiang Li 0001
IEEE Trans. Circuits Syst. Video Technol.3
2014 mCENTRIST: A Multi-Channel Feature Generation Mechanism for Scene Categorization
abstract
mCENTRIST, a new multichannel feature generation mechanism for recognizing scene categories, is proposed in this paper. mCENTRIST explicitly captures the image properties that are encoded jointly by two image channels, which is different from popular multichannel descriptors. In order to avoid the curse of dimensionality, tradeoffs at both feature and channel levels have been executed to make mCENTRIST computationally practical. As a result, mCENTRIST is both efficient and easy to implement. In addition, a hyperopponent color space is proposed by embedding Sobel information into the opponent color space for further performance improvements. Experiments show that mCENTRIST outperforms established multichannel descriptors on four RGB and RGB-near infrared data sets, including aerial orthoimagery, indoor, and outdoor scene category recognition tasks. Experiments also verify that the hyper opponent color space enhances descriptors' performance effectively.
Yang Xiao 0007, Jianxin Wu 0001, Junsong Yuan 0001
IEEE Trans. Image Process.2
2014 Flexible Image Similarity Computation Using Hyper-Spatial Matching
abstract
Spatial pyramid matching (SPM) has been widely used to compute the similarity of two images in computer vision and image processing. While comparing images, SPM implicitly assumes that: in two images from the same category, similar objects will appear in similar locations. However, this is not always the case. In this paper, we propose hyper-spatial matching (HSM), a more flexible image similarity computing method, to alleviate the mis-matching problem in SPM. Besides the match between corresponding regions, HSM considers the relationship of all spatial pairs in two images, which includes more meaningful match than SPM. We propose two learning strategies to learn SVM models with the proposed HSM kernel in image classification, which are hundreds of times faster than a general purpose SVM solver applied to the HSM kernel (in both training and testing). We compare HSM and SPM on several challenging benchmarks, and show that HSM is better than SPM in describing image similarity.
Yu Zhang 0004, Jianxin Wu 0001, Jianfei Cai 0001, Weiyao Lin
IEEE Trans. Image Process.2
2013 Reduced Heteroscedasticity Linear Regression for Nyström Approximation
Hao Yang 0033, Jianxin Wu 0001
IJCAI2
2013 Salient object cutout using Google images
abstract
Given any image input by users, how to automatically cutout the object-of-interest is a challenging problem due to lack of information of the object-of-interest and the background. Saliency detection techniques are able to provide some rough information about object-of-interest since they highlight high-contrast or high attention regions or pixels. However, the generated saliency map is often noisy and directly applying it for segmentation often leads to erroneous results. Motivated by the recent progress on image co-segmentation and internet image retrieval techniques, in this paper, we propose to use the user input image for segmentation as a query image to Google Images and then employ the top returned Google images to build up the knowledge about the object-of-interest in the user input image. Particularly, we develop a lightweight algorithm to learn the knowledge of the object-of-interest in the retrieved images to enhance the saliency map of the input image. Then, the enhanced saliency map is used to initialize the graph-cut to extract the object-of-interest. Experiments with the Mcgill dataset and multiple challenge cases demonstrate the effectiveness of our method in terms of producing a clean cutout.
Hongyuan Zhu 0002, Jianfei Cai 0001, Jianmin Zheng, Jianxin Wu 0001, Nadia Magnenat-Thalmann
ISCAS4
2013 Remora: Sensing resource sharing among smartphone-based body sensor networks
abstract
In many body sensor network (BSN) applications, such as activity recognition for assisted living residents or physical fitness assessment of a sports team, users spend a significant amount of time with one another while performing many of the same activities. We exploit this physical proximity with Remora, a smartphone-based Body Sensor Network activity recognition system which shares sensing resources among neighboring BSNs. Compared to other resource sharing approaches, Remora provides both increased accuracy and significant energy savings. To increase classification accuracy, Remora BSNs share sensors by overhearing neighbors' sensor data transmissions. When sharing, fewer on-body sensors are needed to achieve high accuracy, resulting in energy savings by turning off unneeded sensors. To save phone energy, neighboring BSNs share classifiers: only one classifier is active at a time classifying activities for all neighbors. Remora addresses three major challenges of sharing with physical neighbors: 1) Sharing only when the energy benefit outweighs the cost, 2) Finding and utilizing the shared sensors and classifiers which produce the best combination of accuracy improvement and energy savings, and 3) Providing a lightweight and collaborative classification approach, without the use of a backend server, which adapts to the dynamics of available neighbors. In a two week evaluation with 6 subjects, we show that Remora provides up to a 30% accuracy increase while extending phone battery lifetime by over 65%.
Matthew Keally, Gang Zhou 0002, Guoliang Xing, Jianxin Wu 0001
IWQoS4
2013 A New Network-Based Algorithm for Human Group Activity Recognition in Videos
Gaojian Li, Weiyao Lin, Jianxin Wu 0001, Yuanzhe Chen, Hui Wei 0001
MMM (1)4
2013 Probabilistic classifiers with a generalized Gaussian scale mixture prior
Jianxin Wu 0001, Suiping Zhou
Pattern Recognit.2
2013 A Heat-Map-Based Algorithm for Recognizing Group Activities in Videos
abstract
In this paper, a new heat-map-based algorithm is proposed for group activity recognition. The proposed algorithm first models human trajectories as series of heat sources and then applies a thermal diffusion process to create a heat map (HM) for representing the group activities. Based on this HM, a new key-point-based (KPB) method is used for handling the alignments among HMs with different scales and rotations. A surface-fitting (SF) method is also proposed for recognizing group activities. Our proposed HM feature can efficiently embed the temporal motion information of the group activities while the proposed KPB and SF methods can effectively utilize the characteristics of the HM for activity recognition. Section IV demonstrates the effectiveness of our proposed algorithms.
Weiyao Lin, Hang Chu, Jianxin Wu 0001, Bin Sheng 0001, Zhenzhong Chen 0001
IEEE Trans. Circuits Syst. Video Technol.3
2013 ${\rm C}^{4}$: A Real-Time Object Detection Framework
abstract
A real-time and accurate object detection framework, C(4), is proposed in this paper. C(4) achieves 20 fps speed and the state-of-the-art detection accuracy, using only one processing thread without resorting to special hardware such as GPU. The real-time accurate object detection is made possible by two contributions. First, we conjecture (with supporting experiments) that contour is what we should capture and signs of comparisons among neighboring pixels are the key information to capture contour cues. Second, we show that the CENTRIST visual descriptor is suitable for contour based object detection, because it encodes the sign information and can implicitly represent the global contour. When CENTRIST and linear classifier are used, we propose a computational method that does not need to explicitly generate feature vectors. It involves no image preprocessing or feature vector normalization, and only requires O(1) steps to test an image patch. C(4) is also friendly to further hardware acceleration. It has been applied to detect objects such as pedestrians, faces, and cars on benchmark data sets. It has comparable detection accuracy with state-of-the-art methods, and has a clear advantage in detection speed.
Jianxin Wu 0001, Nini Liu, Christopher Geyer, James M. Rehg
IEEE Trans. Image Process.1
2012 Object Templates for Visual Place Categorization
Hao Yang 0033, Jianxin Wu 0001
ACCV (4)2
2012 Exclusive Visual Descriptor Quantization
Yu Zhang 0004, Jianxin Wu 0001, Weiyao Lin
ACCV (1)2
2012 Power mean SVM for large scale visual classification
abstract
PmSVM (Power Mean SVM), a classifier that trains significantly faster than state-of-the-art linear and non-linear SVM solvers in large scale visual classification tasks, is presented. PmSVM also achieves higher accuracies. A scalable learning method for large vision problems, e.g., with millions of examples or dimensions, is a key component in many current vision systems. Recent progresses have enabled linear classifiers to efficiently process such large scale problems. Linear classifiers, however, usually have inferior accuracies in vision tasks. Non-linear classifiers, on the other hand, may take weeks or even years to train. We propose a power mean kernel and present an efficient learning algorithm through gradient approximation. The power mean kernel family include as special cases many popular additive kernels. Empirically, PmSVM is up to 5 times faster than LIBLINEAR, and two times faster than state-of-the-art additive kernel classifiers. In terms of accuracy, it outperforms state-of-the-art additive kernel implementations, and has major advantages over linear SVM.
Jianxin Wu 0001
CVPR1
2012 Facial expression mapping based on elastic and muscle-distribution-based models
abstract
In this paper, a new algorithm is proposed for facial expression mapping. The proposed algorithm first introduces a new elastic model to balance the global and local warping effects such that the impacts from facial feature differences between people can be avoided, thus more reasonable geometric warping results can be created. Furthermore, a muscle-distribution-based (MD) model is also proposed. The proposed MD model utilizes the muscle distribution information of the human face to evaluate and strengthen the facial illumination details. By this way, the impacts from human face difference as well as the effects of unsuitable noise filtering can be effectively alleviated. Experimental results show that our proposed algorithm can create obviously better facial expression results than the existing methods.
Weiyao Lin, Bin Sheng 0001, Jianxin Wu 0001, Hongxiang Li 0001
ISCAS4
2012 A new heat-map-based algorithm for human group activity recognition
abstract
In this paper, a new heat-map-based (HMB) algorithm is proposed for human group activity recognition. The proposed algorithm first models people trajectories as series of "heat sources" and then applies a thermal diffusion process to create a heat map (HM) for representing the group activities. Based on this heat map, a new surface-fitting (SF) method is also proposed for recognizing human group activities. Our proposed HM feature can efficiently keep the temporal motion information of the group activities while the proposed SF method can effectively catch the characteristics of the heat map for activity recognition. Experimental results demonstrate the effectiveness of our proposed algorithm.
Hang Chu, Weiyao Lin, Jianxin Wu 0001, Xingtong Zhou, Yuanzhe Chen, Hongxiang Li 0001
ACM Multimedia3
2012 Efficient HIK SVM Learning for Image Classification
abstract
Histograms are used in almost every aspect of image processing and computer vision, from visual descriptors to image representations. Histogram intersection kernel (HIK) and support vector machine (SVM) classifiers are shown to be very effective in dealing with histograms. This paper presents contributions concerning HIK SVM for image classification. First, we propose intersection coordinate descent (ICD), a deterministic and scalable HIK SVM solver. ICD is much faster than, and has similar accuracies to, general purpose SVM solvers and other fast HIK SVM training methods. We also extend ICD to the efficient training of a broader family of kernels. Second, we show an important empirical observation that ICD is not sensitive to the C parameter in SVM, and we provide some theoretical analyses to explain this observation. ICD achieves high accuracies in many problems, using its default parameters. This is an attractive property for practitioners, because many image processing tasks are too large to choose SVM parameters using cross-validation.
Jianxin Wu 0001
IEEE Trans. Image Process.1
2011 Real-time human detection using contour cues
abstract
A real-time and accurate human detector, C4, is proposed in this paper. C4achieves 20 fps speed and state-of-the-art detection accuracy, using only one processing thread without resorting to special hardwares like GPU. Real-time accurate human detection is made possible by two contributions. First, we show that contour is exactly what we should capture and signs of comparisons among neighboring pixels are the key information to capture contours. Second, we show that the CENTRIST visual descriptor is particularly suitable for human detection, because it encodes the sign information and can implicitly represent the global contour. When CENTRIST and linear classifier are used, we propose a computational method that does not need to explicitly generate feature vectors. It involves no image pre-processing or feature vector normalization, and only requires O(1) steps to test an image patch. C4is also friendly to further hardware acceleration. In a robot with embedded 1.2GHz CPU, we also achieved accurate and 20 fps high speed human detection.
Jianxin Wu 0001, Christopher Geyer, James M. Rehg
ICRA1
2011 Probit Classifiers with a Generalized Gaussian Scale Mixture Prior
Jianxin Wu 0001, Suiping Zhou
IJCAI2
2011 Exploiting sensing diversity for confident sensing in wireless sensor networks
abstract
Wireless sensor networks for human health monitoring, military surveillance, and disaster warning all have stringent accuracy requirements for detecting or classifying events while maximizing system lifetime. We define meeting such user accuracy requirements as confident sensing. To perform confident sensing and reduce energy, we must address sensing diversity: sensing capability differences among heterogeneous and homogeneous sensors in a specific deployment. We are among the first to explore the impact of sensing diversity on sensor collaboration, exploit diversity for sensing confidence, and apply diversity exploitation for confident sensing coverage. We show that our diversity-exploiting confident coverage problem is NP-hard for any specific deployment and present a practical solution, Wolfpack. Through a distributed and iterative sensor collaboration approach, Wolfpack maximizes a specific deployment's capability to meet user detection requirements and save energy by powering off unneeded nodes. Using real vehicle detection trace data, we demonstrate that Wolfpack provides confident event detection coverage for 30% more detection locations, using 20% less energy than a state of the art approach.
Matthew Keally, Gang Zhou 0002, Guoliang Xing, Jianxin Wu 0001
INFOCOM4
2011 Balance Support Vector Machines Locally Using the Structural Similarity Kernel
Jianxin Wu 0001
PAKDD (1)1
2011 PBN: towards practical activity recognition using smartphone-based body sensor networks
abstract
The vast array of small wireless sensors is a boon to body sensor network applications, especially in the context awareness and activity recognition arena. However, most activity recognition deployments and applications are challenged to provide personal control and practical functionality for everyday use. We argue that activity recognition for mobile devices must meet several goals in order to provide a practical solution: user friendly hardware and software, accurate and efficient classification, and reduced reliance on ground truth. To meet these challenges, we present PBN: Practical Body Networking. Through the unification of TinyOS motes and Android smartphones, we combine the sensing power of on-body wireless sensors with the additional sensing power, computational resources, and user-friendly interface of an Android smartphone. We provide an accurate and efficient classification approach through the use of ensemble learning. We explore the properties of different sensors and sensor data to further improve classification efficiency and reduce reliance on user annotated ground truth. We evaluate our PBN system with multiple subjects over a two week period and demonstrate that the system is easy to use, accurate, and appropriate for mobile devices.
Matthew Keally, Gang Zhou 0002, Guoliang Xing, Jianxin Wu 0001, Andrew J. Pyles
SenSys4
2011 Efficient and Effective Visual Codebook Generation Using Additive Kernels
Jianxin Wu 0001, Wei-Chian Tan, James M. Rehg
J. Mach. Learn. Res.1
2011 CENTRIST: A Visual Descriptor for Scene Categorization
abstract
CENsus TRansform hISTogram (CENTRIST), a new visual descriptor for recognizing topological places or scene categories, is introduced in this paper. We show that place and scene recognition, especially for indoor environments, require its visual descriptor to possess properties that are different from other vision domains (e.g., object recognition). CENTRIST satisfies these properties and suits the place and scene recognition task. It is a holistic representation and has strong generalizability for category recognition. CENTRIST mainly encodes the structural properties within an image and suppresses detailed textural information. Our experiments demonstrate that CENTRIST outperforms the current state of the art in several place and scene recognition data sets, compared with other descriptors such as SIFT and Gist. Besides, it is easy to implement and evaluates extremely fast.
Jianxin Wu 0001, James M. Rehg
IEEE Trans. Pattern Anal. Mach. Intell.1
2010 A Fast Dual Method for HIK SVM Learning
Jianxin Wu 0001
ECCV (2)1
2010 Corrigendum to "Ensembling neural networks: Many could be better than all" [Artificial Intelligence 137 (1-2) (2002) 239-263]
Zhi-Hua Zhou, Jianxin Wu 0001
Artif. Intell.2
2009 Beyond the Euclidean distance: Creating effective visual codebooks using the Histogram Intersection Kernel
abstract
Common visual codebook generation methods used in a Bag of Visual words model, e.g. k-means or Gaussian Mixture Model, use the Euclidean distance to cluster features into visual code words. However, most popular visual descriptors are histograms of image measurements. It has been shown that the Histogram Intersection Kernel (HIK) is more effective than the Euclidean distance in supervised learning tasks with histogram features. In this paper, we demonstrate that HIK can also be used in an unsupervised manner to significantly improve the generation of visual codebooks. We propose a histogram kernel k-means algorithm which is easy to implement and runs almost as fast as k-means. The HIK codebook has consistently higher recognition accuracy over k-means codebooks by 2-4%. In addition, we propose a one-class SVM formulation to create more effective visual code words which can achieve even higher accuracy. The proposed method has established new state-of-the-art performance numbers for 3 popular benchmark datasets on object and scene recognition. In addition, we show that the standard k-median clustering method can be used for visual codebook generation and can act as a compromise between HIK and k-means approaches.
Jianxin Wu 0001, James M. Rehg
ICCV1
2009 Visual Place Categorization: Problem, dataset, and algorithm
abstract
In this paper we describe the problem of visual place categorization (VPC) for mobile robotics, which involves predicting the semantic category of a place from image measurements acquired from an autonomous platform. For example, a robot in an unfamiliar home environment should be able to recognize the functionality of the rooms it visits, such as kitchen, living room, etc. We describe an approach to VPC based on sequential processing of images acquired with a conventional video camera. We identify two key challenges: Dealing with non-characteristic views and integrating restricted-FOV imagery into a holistic prediction. We present a solution to VPC based upon a recently-developed visual feature known as CENTRIST (census transform histogram). We describe a new dataset for VPC which we have recently collected and are making publicly available. We believe this is the first significant, realistic dataset for the VPC problem. It contains the interiors of six different homes with ground truth labels. We use this dataset to validate our solution approach, achieving promising results.
Jianxin Wu 0001, Henrik I. Christensen, James M. Rehg
IROS1
2009 Exploratory Undersampling for Class-Imbalance Learning
abstract
Undersampling is a popular method in dealing with class-imbalance problems, which uses only a subset of the majority class and thus is very efficient. The main deficiency is that many majority class examples are ignored. We propose two algorithms to overcome this deficiency. EasyEnsemble samples several subsets from the majority class, trains a learner using each of them, and combines the outputs of those learners. BalanceCascade trains the learners sequentially, where in each step, the majority class examples that are correctly classified by the current trained learners are removed from further consideration. Experimental results show that both methods have higher Area Under the ROC Curve, F-measure, and G-mean values than many existing class-imbalance learning methods. Moreover, they have approximately the same training time as that of undersampling when the same number of weak classifiers is used, which is significantly faster than other methods.
Xu-Ying Liu, Jianxin Wu 0001, Zhi-Hua Zhou
IEEE Trans. Syst. Man Cybern. Part B2
2008 Where am I: Place instance and category recognition using spatial PACT
abstract
We introduce spatial PACT (Principal component Analysis of Census Transform histograms), a new representation for recognizing instances and categories of places or scenes. Both place instance recognition (“I am in Room 113”) and category recognition (“I am in an office”) have been widely researched. Features that have different discriminative power/invariance tradeoff have been used separately for the two tasks. PACT captures local structures of an image through the Census Transform (CT), while large-scale structures are captured by the strong correlation between neighboring CT values and the histogram. The PCA operation ignores noise in the histogram distribution, computes important “primitive shapes”, and results in a compact representation. Spatial PACT, a spatial pyramid of PACT, further incorporates global structures in the image. Our experiments demonstrate that spatial PACT outperforms the current state-of-the-art in several place and scene recognition, and shape matching datasets. Besides, spatial PACT is easy to implement. It has nearly no parameter to tune, and evaluates extremely fast.
Jianxin Wu 0001, James M. Rehg
CVPR1
2008 On the Design of Cascades of Boosted Ensembles for Face Detection
S. Charles Brubaker, Jianxin Wu 0001, Jie Sun 0004, Matthew D. Mullin, James M. Rehg
Int. J. Comput. Vis.2
2008 Fast Asymmetric Learning for Cascade Face Detection
abstract
A cascade face detector uses a sequence of node classifiers to distinguish faces from non-faces. This paper presents a new approach to design node classifiers in the cascade detector. Previous methods used machine learning algorithms that simultaneously select features and form ensemble classifiers. We argue that if these two parts are decoupled, we have the freedom to design a classifier that explicitly addresses the difficulties caused by the asymmetric learning goal. There are three contributions in this paper. The first is a categorization of asymmetries in the learning goal, and why they make face detection hard. The second is the Forward Feature Selection (FFS) algorithm and a fast pre- omputing strategy for AdaBoost. FFS and the fast AdaBoost can reduce the training time by approximately 100 and 50 times, in comparison to a naive implementation of the AdaBoost feature selection method. The last contribution is Linear Asymmetric Classifier (LAC), a classifier that explicitly handles the asymmetric learning goal as a well-defined constrained optimization problem. We demonstrated experimentally that LAC results in improved ensemble classifier performance.
Jianxin Wu 0001, S. Charles Brubaker, Matthew D. Mullin, James M. Rehg
IEEE Trans. Pattern Anal. Mach. Intell.1
2007 A Scalable Approach to Activity Recognition based on Object Use
abstract
We propose an approach to activity recognition based on detecting and analyzing the sequence of objects that are being manipulated by the user. In domains such as cooking, where many activities involve similar actions, object-use information can be a valuable cue. In order for this approach to scale to many activities and objects, however, it is necessary to minimize the amount of human-labeled data that is required for modeling. We describe a method for automatically acquiring object models from video without any explicit human supervision. Our approach leverages sparse and noisy readings from RFID tagged objects, along with common-sense knowledge about which objects are likely to be used during a given activity, to bootstrap the learning process. We present a dynamic Bayesian network model which combines RFID and video data to jointly infer the most likely activity and object labels. We demonstrate that our approach can achieve activity recognition rates of more than 80% on a real-world dataset consisting of 16 household activities involving 33 objects with significant background clutter. We show that the combination of visual object recognition with RFID data is significantly more effective than the RFID sensor alone. Our work demonstrates that it is possible to automatically learn object models from video of household activities and employ these models for activity recognition, without requiring any explicit human labeling.
Jianxin Wu 0001, Adebola Osuntogun, Tanzeem Choudhury, Matthai Philipose, James M. Rehg
ICCV1
2006 Exploratory Under-Sampling for Class-Imbalance Learning
abstract
Under-sampling is a class-imbalance learning method which uses only a subset of major class examples and thus is very efficient. The main deficiency is that many major class examples are ignored. We propose two algorithms to overcome the deficiency. EasyEnsemble samples several subsets from the major class, trains a learner using each of them, and combines the outputs of those learners. BalanceCascade is similar to EasyEnsemble except that it removes correctly classified major class examples of trained learners from further consideration. Experiments show that both of the proposed algorithms have better AUC scores than many existing class-imbalance learning methods. Moreover, they have approximately the same training time as that of under-sampling, which trains significantly faster than other methods.
Xu-Ying Liu, Jianxin Wu 0001, Zhi-Hua Zhou
ICDM2
2005 Linear Asymmetric Classifier for cascade detectors
abstract
The detection of faces in images is fundamentally a rare event detection problem. Cascade classifiers provide an efficient computational solution, by leveraging the asymmetry in the distribution of faces vs. non-faces. Training a cascade classifier in turn requires a solution for the following subproblems: Design a classifier for each node in the cascade with very high detection rate but only moderate false positive rate. While there are a few strategies in the literature for indirectly addressing this asymmetric node learning goal, none of them are based on a satisfactory theoretical framework. We present a mathematical characterization of the node-learning problem and describe an effective closed form approximation to the optimal solution, which we call the Linear Asymmetric Classifier (LAC). We first use AdaBoost or AsymBoost to select features, and use LAC to learn a linear discriminant function to achieve the node learning goal. Experimental results on face detection show that LAC can improve the detection performance in comparison to standard methods. We also show that Fisher Discriminant Analysis on the features selected by AdaBoost yields better performance than AdaBoost itself.
Jianxin Wu 0001, Matthew D. Mullin, James M. Rehg
ICML1
2003 Learning a Rare Event Detection Cascade by Direct Feature Selection
abstract
Face detection is a canonical example of a rare event detection prob- lem, in which target patterns occur with much lower frequency than non- targets. Out of millions of face-sized windows in an input image, for ex- ample, only a few will typically contain a face. Viola and Jones recently proposed a cascade architecture for face detection which successfully ad- dresses the rare event nature of the task. A central part of their method is a feature selection algorithm based on AdaBoost. We present a novel cascade learning algorithm based on forward feature selection which is two orders of magnitude faster than the Viola-Jones approach and yields classifiers of equivalent quality. This faster method could be used for more demanding classification tasks, such as on-line learning.
Jianxin Wu 0001, James M. Rehg, Matthew D. Mullin
NIPS1
2003 Efficient face candidates selector for face detection
Jianxin Wu 0001, Zhi-Hua Zhou
Pattern Recognit.1
2002 Ensembling neural networks: Many could be better than all
Zhi-Hua Zhou, Jianxin Wu 0001
Artif. Intell.2
2002 Face recognition with one training image per person
Jianxin Wu 0001, Zhi-Hua Zhou
Pattern Recognit. Lett.1
2001 Genetic Algorithm based Selective Neural Network Ensemble
Zhi-Hua Zhou, Jianxin Wu 0001, Yuan Jiang 0001, Shifu Chen
IJCAI2
2001 Combining Regression Estimators: GA-Based Selective Neural Network Ensemble
abstract
Neural network ensemble is a learning paradigm where a collection of neural networks is trained for the same task. In this paper, the relationship between the generalization ability of the neural network ensemble and the correlation of the individual neural networks constituting the ensemble is analyzed in the context of combining neural regression estimators, which reveals that ensembling a selective subset of trained networks is superior to ensembling all the trained networks in some cases. Based on such recognition, an approach named GASEN is proposed. GASEN trains a number of individual neural networks at first. Then it assigns random weights to the individual networks and employs a genetic algorithm to evolve those weights so that they can characterize to some extent the importance of the individual networks in constituting an ensemble. Finally it selects an optimum subset of individual networks based on the evolved weights to make up the ensemble. Experimental results show that, comparing with a popular ensemble approach, i.e., averaging all, and a theoretically optimum selective ensemble approach, i.e. enumerating, GASEN has preferable performance in generating ensembles with strong generalization ability in relatively small computational cost. This paper also analyzes the working mechanism of GASEN from the view of error-ambiguity decomposition, which reveals that GASEN improves generalization ability mainly through reducing the average generalization error of the individual neural networks constituting the ensemble.
Zhi-Hua Zhou, Jianxin Wu 0001, Zhaoqian Chen
Int. J. Comput. Intell. Appl.2