EDBT 2026 Demo / reviewers in the wild / expert
Fang Wan 0001
dblp:01/845-1
· DBLP profile ↗
38ranked-venue papers
6as first author
26since 2021 · last 2026
0000-0002-8083-9257ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 32 · 5 first-author · 23 since 2021Graphics, computer vision, multimedia, augmented reality and games · 24 · 4 first-author · 15 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Compactness driven Co-learning for crowd counting and localization
Ziheng Yan, Xinyan Liu 0008, Guorong Li, Weigang Zhang, Fang Wan 0001, Qingming Huang |
Pattern Recognit. | 5 |
| 2025 | Timestep Embedding Tells: It's Time to Cache for Video Diffusion ModelabstractAs a fundamental backbone for video generation, diffusion models are challenged by low inference speed due to the sequential nature of denoising. Previous methods speed up the models by caching and reusing model outputs at uniformly selected timesteps. However, such a strategy neglects the fact that differences among model outputs are not uniform across timesteps, which hinders selecting the appropriate model outputs to cache, leading to a poor balance between inference efficiency and visual quality. In this study, we introduce Timestep Embedding Aware Cache (TeaCache), a training-free caching approach that estimates and leverages the fluctuating differences among model outputs across timesteps. Rather than directly using the time-consuming model outputs, TeaCache focuses on model inputs, which have a strong correlation with the modeloutputs while incurring negligible computational cost. TeaCache first modulates the noisy inputs using the timestep embeddings to ensure their differences better approximating those of model outputs. TeaCache then introduces a rescaling strategy to refine the estimated differences and utilizes them to indicate output caching. Experiments show that TeaCache achieves up to 4.41× acceleration over Open-Sora-Plan with negligible (-0.07% Vbench score) degradation of visual quality. Feng Liu 0050, Shiwei Zhang 0001, Yujie Wei 0001, Haonan Qiu, Yuzhong Zhao, Yingya Zhang, Qixiang Ye, Fang Wan 0001 |
CVPR | 9 |
| 2025 | DynRefer: Delving into Region-level Multimodal Tasks via Dynamic ResolutionabstractOne fundamental task of multimodal models is to translate referred image regions to human preferred language descriptions. Existing methods, however, ignore the resolution adaptability needs of different tasks, which hinders them to find out precise language descriptions. In this study, we propose a DynRefer approach, to pursue high-accuracy region-level referring through mimicking the resolution adaptability of human visual cognition. During training, DynRefer stochastically aligns language descriptions of multimodal tasks with images of multiple resolutions, which are constructed by nesting a set of random views around the referred region. During inference, DynRefer performs selectively multimodal referring by sampling proper region representations for tasks from the nested views based on image and task priors. This allows the visual information for referring to better match human preferences, thereby improving the representational adaptability of region-level multimodal models. Experiments show that DynRefer brings mutual improvement upon broad tasks including region-level captioning, open-vocabulary region recognition and attribute detection. Furthermore, DynRefer achieves state-of-the-art results on multiple region-level multimodal tasks using a single model. Code is available at https://github.com/callsys/DynRefer. Yuzhong Zhao, Feng Liu 0050, Mingxiang Liao, Chen Gong 0005, Qixiang Ye, Fang Wan 0001 |
CVPR | 7 |
| 2025 | Dual Discrepancy-Based Continuation Learning for Hybrid Supervised Multi-View 3D Object Detection
Tianyu Wang 0028, Feng Liu 0050, Jianbin Jiao, Fang Wan 0001 |
PRCV (17) | 5 |
| 2025 | Discriminatively Matched Part Tokens for Pointly Supervised Instance Segmentation
Zonghao Guo, Fang Wan 0001, Mingxiang Liao, Qixiang Ye |
Int. J. Comput. Vis. | 2 |
| 2025 | Hierarchical AttentionShift for Pointly Supervised Instance SegmentationabstractPointly supervised instance segmentation (PSIS) remains a challenging task when appearance variances across object parts cause semantic inconsistency. In this article, we propose a hierarchical AttentionShift approach, to solve the semantic inconsistency issue through exploiting the hierarchical nature of semantics and the flexibility of key-point representation. The estimation of hierarchical attention is defined upon key-point sets. The representative key points are iteratively estimated spatially and in the feature space to capture the fine-grained semantics and cover the full object extent. Hierarchical AttentionShift is performed at instance, part, and fine-grained levels, optimizing object semantics while promoting the conventional self-attention activation to hierarchical activation with local refinement. Experiments on PASCAL VOC 2012 Aug and MS-COCO 2017 benchmarks show that hierarchical AttentionShift improves the state-of-the-art (SOTA) method by 10.4% and 7.0% upon mean average precision (mAP)50, respectively. When applying hierarchical AttentionShift to the segment anything model (SAM), 9.4% AP improvement on the COCO test-dev is achieved. Hierarchical AttentionShift provides a fresh insight to regularize the self-attention mechanism for fine-grained vision tasks. The code is available at github.com/MingXiangL/AttentionShift. Mingxiang Liao, Fang Wan 0001, Zonghao Guo, Qixiang Ye |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | Ray Denoising: Depth-Aware Hard Negative Sampling for Multi-view 3D Object Detection
Feng Liu 0050, Tengteng Huang, Qianjing Zhang, Fang Wan 0001, Qixiang Ye, Yanzhao Zhou |
ECCV (49) | 6 |
| 2024 | ControlCap: Controllable Region-Level Captioning
Yuzhong Zhao, Zonghao Guo, Weijia Wu 0001, Chen Gong 0005, Qixiang Ye, Fang Wan 0001 |
ECCV (38) | 7 |
| 2024 | Evaluation of Text-to-Video Generation Models: A Dynamics PerspectiveabstractComprehensive and constructive evaluation protocols play an important role when developing sophisticated text-to-video (T2V) generation models. Existing evaluation protocols primarily focus on temporal consistency and content continuity, yet largely ignore dynamics of video content. Such dynamics is an essential dimension measuring the visual vividness and the honesty of video content to text prompts. In this study, we propose an effective evaluation protocol, termed DEVIL, which centers on the dynamics dimension to evaluate T2V generation models, as well as improving existing evaluation metrics. In practice, we define a set of dynamics scores corresponding to multiple temporal granularities, and a new benchmark of text prompts under multiple dynamics grades. Upon the text prompt benchmark, we assess the generation capacity of T2V models, characterized by metrics of dynamics ranges and T2V alignment. Moreover, we analyze the relevance of existing metrics to dynamics metrics, improving them from the perspective of dynamics. Experiments show that DEVIL evaluation metrics enjoy up to about 90\% consistency with human ratings, demonstrating the potential to advance T2V generation models. Mingxiang Liao, Hannan Lu, Qixiang Ye, Wangmeng Zuo, Fang Wan 0001, Tianyu Wang 0028, Yuzhong Zhao, Jingdong Wang 0001, Xinyu Zhang 0017 |
NeurIPS | 5 |
| 2024 | TS-CAM: Token Semantic Coupled Attention Map for Weakly Supervised Object LocalizationabstractWeakly supervised object localization (WSOL), which trains object localization models using solely image category annotations, remains a challenging problem. Existing approaches based on convolutional neural networks (CNNs) tend to miss full object extent while activating discriminative object parts. Based on our analysis, this is caused by CNN's intrinsic characteristics, which experiences difficulty to capture object semantics at long distances. In this article, we introduce the vision transformer to WSOL, with the aim to capture long-range semantic dependency of features by leveraging transformer's cascaded self-attention mechanism. We propose the token semantic coupled attention map (TS-CAM) method, which first decomposes class-aware semantics and then couples the semantics with attention maps for semantic-aware activation. To capture object semantics at long distances and avoid partial activation, TS-CAM performs spatial embedding by partitioning an image to a set of patch tokens. To incorporate object category information to patch tokens, TS-CAM reallocates category-related semantics to each patch token. The patch tokens are finally coupled with attention maps which are semantic-agnostic to perform semantic-aware object localization. By introducing semantic tokens to produce semantic-aware attention maps, we further explore the capability of TS-CAM for multicategory object localization. Experiments show that TS-CAM outperforms its CNN-CAM counterpart by 11.6% and 28.9% on ILSVRC and CUB-200-2011 datasets, respectively, improving the state-of-the-art with large margins. TS-CAM also demonstrates superiority for multicategory object localization on the Pascal VOC dataset. The code is available at github.com/yuanyao366/ts-cam-extension. Fang Wan 0001, Wei Gao 0050, Xingjia Pan, Zhiliang Peng, Qi Tian 0001, Qixiang Ye |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | AttentionShift: Iteratively Estimated Part-Based Attention Map for Pointly Supervised Instance SegmentationabstractPointly supervised instance segmentation (PSIS) learns to segment objects using a single point within the object extent as supervision. Challenged by the non-negligible semantic variance between object parts, however, the single supervision point causes semantic bias and false segmentation. In this study, we propose an AttentionShift method, to solve the semantic bias issue by iteratively decomposing the instance attention map to parts and estimating fine-grained semantics of each part. AttentionShift consists of two modules plugged on the vision transformer backbone: (i) token querying for pointly supervised attention map generation, and (ii) key-point shift, which re-estimates part-based attention maps by key-point filtering in the feature space. These two steps are iteratively performed so that the part-based attention maps are optimized spatially as well as in the feature space to cover full object extent. Experiments on PASCAL VOC and MS COCO 2017 datasets show that AttentionShift respectively improves the state-of-the-art of by 7.7% and 4.8% under [email protected], setting a solid PSIS baseline using vision transformer. Mingxiang Liao, Zonghao Guo, Yuze Wang 0004, Bailan Feng, Fang Wan 0001 |
CVPR | 6 |
| 2023 | Integrally Migrating Pre-trained Transformer Encoder-decoders for Visual Object DetectionabstractModern object detectors have taken the advantages of backbone networks pre-trained on large scale datasets. Except for the backbone networks, however, other components such as the detector head and the feature pyramid network (FPN) remain trained from scratch, which hinders the generalization capacity of detectors. In this study, we propose to integrally migrate pre-trained transformer encoder-decoders (imTED) to a detector, constructing a feature extraction path which is "fully pre-trained" so that detectors’ generalization capacity is maximized. The essential differences between imTED with the baseline detector are twofold: (1) migrating the pre-trained transformer decoder to the detector head while removing the randomly initialized FPN from the feature extraction path; and (2) defining a multi-scale feature modulator (MFM) to enhance scale adaptability. Such designs not only reduce randomly initialized parameters significantly but also unify detector training with representation learning intendedly. Experiments on the MS COCO object detection dataset show that imTED consistently outperforms its counterparts by ~2.4 AP. Without bells and whistles, imTED improves the state-of-the-art of few-shot object detection by up to 7.6 AP. Code is released at https://github.com/LiewFeng/imTED. Feng Liu 0050, Xiaosong Zhang 0004, Zhiliang Peng, Zonghao Guo, Fang Wan 0001, Xiangyang Ji, Qixiang Ye |
ICCV | 5 |
| 2023 | Generative Prompt Model for Weakly Supervised Object LocalizationabstractWeakly supervised object localization (WSOL) remains challenging when learning object localization models from image category labels. Conventional methods that discriminatively train activation models ignore representative yet less discriminative object parts. In this study, we propose a generative prompt model (GenPromp), defining the first generative pipeline to localize less discriminative object parts by formulating WSOL as a conditional image denoising procedure. During training, GenPromp converts image category labels to learnable prompt embeddings which are fed to a generative model to conditionally recover the input image with noise and learn representative embeddings. During inference, GenPromp combines the representative embeddings with discriminative embeddings (queried from an off-the-shelf vision-language model) for both representative and discriminative capacity. The combined embeddings are finally used to generate multi-scale high-quality attention maps, which facilitate localizing full object extent. Experiments on CUB-200-2011 and ILSVRC show that GenPromp respectively outperforms the best discriminative models by 5.2% and 5.6% (Top-1 Loc), setting a solid baseline for WSOL with the generative model. Code is available at https://github.com/callsys/GenPromp. Yuzhong Zhao, Qixiang Ye, Weijia Wu 0001, Chunhua Shen, Fang Wan 0001 |
ICCV | 5 |
| 2023 | Multiple Instance Differentiation Learning for Active Object DetectionabstractDespite the substantial progress of active learning for image recognition, there lacks a systematic investigation of instance-level active learning for object detection. In this paper, we propose to unify instance uncertainty calculation with image uncertainty estimation for informative image selection, creating a multiple instance differentiation learning (MIDL) method for instance-level active learning. MIDL consists of a classifier prediction differentiation module and a multiple instance differentiation module. The former leverages two adversarial instance classifiers trained on the labeled and unlabeled sets to estimate instance uncertainty of the unlabeled set. The latter treats unlabeled images as instance bags and re-estimates image-instance uncertainty using the instance classification model in a multiple instance learning fashion. Through weighting the instance uncertainty using instance class probability and instance objectness probability under the total probability formula, MIDL unifies the image uncertainty with instance uncertainty in the Bayesian theory framework. Extensive experiments validate that MIDL sets a solid baseline for instance-level active learning. On commonly used object detection datasets, it outperforms other state-of-the-art methods by significant margins, particularly when the labeled sets are small. Fang Wan 0001, Qixiang Ye, Tianning Yuan, Songcen Xu, Jianzhuang Liu, Xiangyang Ji, Qingming Huang |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | End-to-End Weakly Supervised Object Detection with Sparse Proposal Evolution
Mingxiang Liao, Fang Wan 0001, Zhenjun Han, Jialing Zou, Yuze Wang 0004, Bailan Feng, Qixiang Ye |
ECCV (9) | 2 |
| 2022 | Learning to Match Anchors for Visual Object DetectionabstractModern CNN-based object detectors assign anchors for ground-truth objects under the restriction of object-anchor Intersection-over-Union (IoU). In this study, we propose a learning-to-match (LTM) method to break IoU restriction, allowing objects to match anchors in a flexible manner. LTM updates hand-crafted anchor assignment to "free" anchor matching by formulating detector training in the Maximum Likelihood Estimation (MLE) framework. During the training phase, LTM is implemented by converting the detection likelihood to anchor matching loss functions which are plug-and-play. Minimizing the matching loss functions drives learning and selecting features which best explain a class of objects with respect to both classification and localization. LTM is extended from anchor-based detectors to anchor-free detectors, validating the general applicability of learnable object-feature matching mechanism for visual object detection. Experiments on MS COCO dataset demonstrate that LTM detectors consistently outperform counterpart detectors with significant margins. The last but not the least, LTM requires negligible computational cost in both training and inference phases as it does not involve any additional architecture or parameter. Code has been made publicly available. Xiaosong Zhang 0004, Fang Wan 0001, Chang Liu 0047, Xiangyang Ji, Qixiang Ye |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | Discrepant multiple instance learning for weakly supervised object detection
Wei Gao 0050, Fang Wan 0001, Jun Yue 0004, Songcen Xu, Qixiang Ye |
Pattern Recognit. | 2 |
| 2022 | Domain Contrast for Domain Adaptive Object DetectionabstractDespite of the substantial progress of visual object detection, models trained in one video domain often fail to generalize well to others due to the change of camera configurations, lighting conditions, and object person views. In this paper, we present Domain Contrast (DC), a simple yet effective approach inspired by contrastive learning for training domain adaptive detectors. DC is deduced from the error bound minimization perspective of a transferred model, and is implemented with cross-domain contrast loss which is plug-and-play. By minimizing cross-domain contrast loss, DC transfers detectors across domains while naturally alleviating the class imbalance issue in the target domain. DC can be applied at either image level or region level, consistently improving detectors’ discriminability while maintaining the transferability. Extensive experiments on commonly used benchmarks show that DC improves the baseline and state-of-the-art by significant margins, while demonstrating great potential for large domain divergence. Code is released athttps://github.com/PhoneSix/Domain-Contrast. Feng Liu 0050, Xiaosong Zhang 0004, Fang Wan 0001, Xiangyang Ji, Qixiang Ye |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Adversarial Prototype Learning for Hyperspectral Image ClassificationabstractIn hyperspectral image (HSI) classification, the training set often contains a very limited number of high-dimensional samples, which can cause overfitting problems, especially in deep learning (DL) frameworks. This situation worsens when a bias exists between the feature distributions of the training and testing sets. In this article, we propose a novel method, referred to as adversarial prototype learning (APL), for learning an accurate HSI classification model in a uniform manner when the training set contains few, high-dimensional, and biased samples. APL consists of a prototype learning module (PLM) and an adversarial alignment module (AAM). The PLM aims to alleviate overfitting by training prototypical classifiers with a simple inductive bias in the initial feature space. The AAM aims to reduce the bias between the feature distributions of the training and testing sets using two adversarial prototypical classifiers learned by the PLM. Iteratively training the PLM and AAM results in alignment of the feature distributions between the training and testing sets while improving the generalization ability of the prototypical classifiers. The theoretical analysis indicates that APL is able to lower the upper error bound when classifying testing samples. We further apply APL in a DL framework to establish the adversarial prototypical network (APNet) architecture. Experimental results on four publicly available HSI datasets demonstrate that the proposed APNet alleviates overfitting, aligns the feature distributions between the training and testing sets, and achieves state-of-the-art performance. Shuai Wang 0059, Bo Du 0001, Dingwen Zhang, Fang Wan 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Part-Based Semantic Transform for Few-Shot Semantic SegmentationabstractFew-shot semantic segmentation remains an open problem for the lack of an effective method to handle the semantic misalignment between objects. In this article, we propose part-based semantic transform (PST) and target at aligning object semantics in support images with those in query images by semantic decomposition-and-match. The semantic decomposition process is implemented with prototype mixture models (PMMs), which use an expectation-maximization (EM) algorithm to decompose object semantics into multiple prototypes corresponding to object parts. The semantic match between prototypes is performed with a min-cost flow module, which encourages correct correspondence while depressing mismatches between object parts. With semantic decomposition-and-match, PST enforces the network's tolerance to objects' appearance and/or pose variation and facilities channelwise and spatial semantic activation of objects in query images. Extensive experiments on Pascal VOC and MS-COCO datasets show that PST significantly improves upon state-of-the-arts. In particular, on MS-COCO, it improves the performance of five-shot semantic segmentation by up to 7.79% with a moderate cost of inference speed and model size. Code for PST is released at https://github.com/Yang-Bob/PST. Boyu Yang 0002, Fang Wan 0001, Chang Liu 0047, Xiangyang Ji, Qixiang Ye |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2022 | Continuation Multiple Instance Learning for Weakly and Fully Supervised Object DetectionabstractWeakly supervised object detection (WSOD) is a challenging task that requires simultaneously learning object detectors and estimating object locations under the supervision of image category labels. Many WSOD methods that adopt multiple instance learning (MIL) have nonconvex objective functions and, therefore, are prone to get stuck in local minima (falsely localize object parts) while missing full object extent during training. In this article, we introduce classical continuation optimization into MIL, thereby creating continuation MIL (C-MIL) with the aim to alleviate the nonconvexity problem in a systematic way. To fulfill this purpose, we partition instances into class-related and spatially related subsets and approximate MIL's objective function with a series of smoothed objective functions defined within the subsets. We further propose a parametric strategy to implement continuation smooth functions, which enables C-MIL to be applied to instance selection tasks in a uniform manner. Optimizing smoothed loss functions prevents the training procedure from falling prematurely into local minima and facilities learning full object extent. Extensive experiments demonstrate the superiority of CMIL over conventional MIL methods. As a general instance selection method, C-MIL is also applied to supervised object detection to optimize anchors/features, improving the detection performance with a significant margin. Qixiang Ye, Fang Wan 0001, Chang Liu 0047, Qingming Huang, Xiangyang Ji |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2021 | Agreement-Discrepancy-Selection: Active Learning with Progressive Distribution AlignmentabstractIn active learning, the ignorance of aligning unlabeled samples' distribution with that of labeled samples hinders the model trained upon labeled samples from selecting informative unlabeled samples. In this paper, we propose an agreement-discrepancy-selection (ADS) approach, and target at unifying distribution alignment with sample selection by introducing adversarial classifiers to the convolutional neural network (CNN). Minimizing classifiers' prediction discrepancy (maximizing prediction agreement) drives learning CNN features to reduce the distribution bias of labeled and unlabeled samples, while maximizing classifiers' discrepancy highlights informative samples. Iterative optimization of agreement and discrepancy loss calibrated with an entropy function drives aligning sample distributions in a progressive fashion for effective active learning. Experiments on image classification and object detection tasks demonstrate that ADS is task-agnostic, while significantly outperforms the previous methods when the labeled sets are small. Mengying Fu, Tianning Yuan, Fang Wan 0001, Songcen Xu, Qixiang Ye |
AAAI | 3 |
| 2021 | Nearest Neighbor Classifier Embedded Network for Active LearningabstractDeep neural networks (DNNs) have been widely applied to active learning. Despite of its effectiveness, the generalization ability of the discriminative classifier (the softmax classifier) is questionable when there is a significant distribution bias between the labeled set and the unlabeled set. In this paper, we attempt to replace the softmax classifier in deep neural network with a nearest neighbor classifier, considering its progressive generalization ability within the unknown sub-space. Our proposed active learning approach, termed nearest Neighbor Classifier Embedded network (NCE-Net), targets at reducing the risk of over-estimating unlabeled samples while improving the opportunity to query informative samples. NCE-Net is conceptually simple but surprisingly powerful, as justified from the perspective of the subset information, which defines a metric to quantify model generalization ability in active learning. Experimental results show that, with simple selection based on rejection or confusion confidence, NCE-Net improves state-of-the-arts on image classification and object detection tasks with significant margins. Fang Wan 0001, Tianning Yuan, Mengying Fu, Xiangyang Ji, Qingming Huang, Qixiang Ye |
AAAI | 1 |
| 2021 | Strengthen Learning Tolerance for Weakly Supervised Object LocalizationabstractWeakly supervised object localization (WSOL) aims at learning to localize objects of interest by only using the image-level labels as the supervision. While numerous efforts have been made in this field, recent approaches still suffer from two challenges: one is the part domination issue while the other is the learning robustness issue. Specifically, the former makes the localizer prone to the local discriminative object regions rather than the desired whole object, and the latter makes the localizer over-sensitive to the variations of the input images so that one can hardly obtain localization results robust to the arbitrary visual stimulus. To solve these issues, we propose a novel framework to strengthen the learning tolerance, referred to as SLT-Net, for WSOL. Specifically, we consider two-fold learning tolerance strengthening mechanisms. One is the semantic tolerance strengthening mechanism, which allows the localizer to make mistakes for classifying similar semantics so that it will not concentrate too much on the discriminative local regions. The other is the visual stimuli tolerance strengthening mechanism, which enforces the localizer to be robust to different image transformations so that the prediction quality will not be sensitive to each specific input image. Finally, we implement comprehensive experimental comparisons on two widely-used datasets CUB and ILSVRC2012, which demonstrate the effectiveness of our proposed approach. Guangyu Guo 0001, Junwei Han 0001, Fang Wan 0001, Dingwen Zhang |
CVPR | 3 |
| 2021 | Multiple Instance Active Learning for Object DetectionabstractDespite the substantial progress of active learning for image recognition, there still lacks an instance-level active learning method specified for object detection. In this paper, we propose Multiple Instance Active Object Detection (MI-AOD), to select the most informative images for detector training by observing instance-level uncertainty. MI-AOD defines an instance uncertainty learning module, which leverages the discrepancy of two adversarial instance classifiers trained on the labeled set to predict instance uncertainty of the unlabeled set. MI-AOD treats unlabeled images as instance bags and feature anchors in images as instances, and estimates the image uncertainty by re-weighting instances in a multiple instance learning (MIL) fashion. Iterative instance uncertainty learning and re-weighting facilitate suppressing noisy instances, toward bridging the gap between instance uncertainty and image-level uncertainty. Experiments validate that MI-AOD sets a solid baseline for instance-level active learning. On commonly used object detection datasets, MI-AOD outperforms state-of-the-art methods with significant margins, particularly when the labeled sets are small. Code is available at https://github.com/yuantn/MI-AOD. Tianning Yuan, Fang Wan 0001, Mengying Fu, Jianzhuang Liu, Songcen Xu, Xiangyang Ji, Qixiang Ye |
CVPR | 2 |
| 2021 | TS-CAM: Token Semantic Coupled Attention Map for Weakly Supervised Object LocalizationabstractWeakly supervised object localization (WSOL) is a challenging problem when given image category labels but requires to learn object localization models. Optimizing a convolutional neural network (CNN) for classification tends to activate local discriminative regions while ignoring complete object extent, causing the partial activation issue. In this paper, we argue that partial activation is caused by the intrinsic characteristics of CNN, where the convolution operations produce local receptive fields and experience difficulty to capture long-range feature dependency among pixels. We introduce the token semantic coupled attention map (TS-CAM) to take full advantage of the self-attention mechanism in visual transformer for long-range dependency extraction. TS-CAM first splits an image into a sequence of patch tokens for spatial embedding, which produce attention maps of long-range visual dependency to avoid partial activation. TS-CAM then re-allocates category-related semantics for patch tokens, enabling each of them to be aware of object categories. TS-CAM finally couples the patch tokens with the semantic-agnostic attention map to achieve semantic-aware localization. Experiments on the ILSVRC/CUB-200-2011 datasets show that TS-CAM outperforms its CNN-CAM counterparts by 7.1%/27.1% for WSOL, achieving state-of-the-art performance. Code is available at https://github.com/vasgaowei/TS-CAM Wei Gao 0050, Fang Wan 0001, Xingjia Pan, Zhiliang Peng, Qi Tian 0001, Zhenjun Han, Bolei Zhou, Qixiang Ye |
ICCV | 2 |
| 2019 | Orthogonal Decomposition Network for Pixel-Wise Binary ClassificationabstractThe weight sharing scheme and spatial pooling operations in Convolutional Neural Networks (CNNs) introduce semantic correlation to neighboring pixels on feature maps and therefore deteriorate their pixel-wise classification performance. In this paper, we implement an Orthogonal Decomposition Unit (ODU) that transforms a convolutional feature map into orthogonal bases targeting at de-correlating neighboring pixels on convolutional features. In theory, complete orthogonal decomposition produces orthogonal bases which can perfectly reconstruct any binary mask (ground-truth). In practice, we further design incomplete orthogonal decomposition focusing on de-correlating local patches which balances the reconstruction performance and computational cost. Fully Convolutional Networks (FCNs) implemented with ODUs, referred to as Orthogonal Decomposition Networks (ODNs), learn de-correlated and complementary convolutional features and fuse such features in a pixel-wise selective manner. Over pixel-wise binary classification tasks for two-dimensional image processing, specifically skeleton detection, edge detection, and saliency detection, and one-dimensional keypoint detection, specifically S-wave arrival time detection for earthquake localization, ODNs consistently improves the state-of-the-arts with significant margins. Chang Liu 0042, Fang Wan 0001, Wei Ke 0003, Zhuowei Xiao, Xiaosong Zhang 0004, Qixiang Ye |
CVPR | 2 |
| 2019 | SIXray: A Large-Scale Security Inspection X-Ray Benchmark for Prohibited Item Discovery in Overlapping ImagesabstractIn this paper, we present a large-scale dataset and establish a baseline for prohibited item discovery in Security Inspection X-ray images. Our dataset, named SIXray, consists of 1,059,231 X-ray images, in which 6 classes of 8,929 prohibited items are manually annotated. It raises a brand new challenge of overlapping image data, meanwhile shares the same properties with existing datasets, including complex yet meaningless contexts and class imbalance. We propose an approach named class-balanced hierarchical refinement (CHR) to deal with these difficulties. CHR assumes that each input image is sampled from a mixture distribution, and that deep networks require an iterative process to infer image contents accurately. To accelerate, we insert reversed connections to different network backbones, delivering high-level visual cues to assist mid-level features. In addition, a class-balanced loss function is designed to maximally alleviate the noise introduced by easy negative samples. We evaluate CHR on SIXray with different ratios of positive/negative samples. Compared to the baselines, CHR enjoys a better ability of discriminating objects especially using mid-level features, which offers the possibility of using a weakly-supervised approach towards accurate object localization. In particular, the advantage of CHR is more significant in the scenarios with fewer positive training samples, which demonstrates its potential application in real-world security inspection. Caijing Miao, Lingxi Xie, Fang Wan 0001, Chi Su, Hongye Liu, Jianbin Jiao, Qixiang Ye |
CVPR | 3 |
| 2019 | C-MIL: Continuation Multiple Instance Learning for Weakly Supervised Object DetectionabstractWeakly supervised object detection (WSOD) is a challenging task when provided with image category supervision but required to simultaneously learn object locations and object detectors. Many WSOD approaches adopt multiple instance learning (MIL) and have non-convex loss functions which are prone to get stuck into local minima (falsely localize object parts) while missing full object extent during training. In this paper, we introduce a continuation optimization method into MIL and thereby creating continuation multiple instance learning (C-MIL), with the intention of alleviating the non-convexity problem in a systematic way. We partition instances into spatially related and class related subsets, and approximate the original loss function with a series of smoothed loss functions defined within the subsets. Optimizing smoothed loss functions prevents the training procedure falling prematurely into local minima and facilitates the discovery of Stable Semantic Extremal Regions (SSERs) which indicate full object extent. On the PASCAL VOC 2007 and 2012 datasets, C-MIL improves the state-of-the-art of weakly supervised object detection and weakly supervised object localization with large margins. Fang Wan 0001, Chang Liu 0042, Wei Ke 0003, Xiangyang Ji, Jianbin Jiao, Qixiang Ye |
CVPR | 1 |
| 2019 | DANet: Divergent Activation for Weakly Supervised Object LocalizationabstractWeakly supervised object localization remains a challenge when learning object localization models from image category labels. Optimizing image classification tends to activate object parts and ignore the full object extent, while expanding object parts into full object extent could deteriorate the performance of image classification. In this paper, we propose a divergent activation (DA) approach, and target at learning complementary and discriminative visual patterns for image classification and weakly supervised object localization from the perspective of discrepancy. To this end, we design hierarchical divergent activation (HDA), which leverages the semantic discrepancy to spread feature activation, implicitly. We also propose discrepant divergent activation (DDA), which pursues object extent by learning mutually exclusive visual patterns, explicitly. Deep networks implemented with HDA and DDA, referred to as DANets, diverge and fuse discrepant yet discriminative features for image classification and object localization in an end-to-end manner. Experiments validate that DANets advance the performance of object localization while maintaining high performance of image classification on CUB-200 and ILSVRC datasets. Haolan Xue, Chang Liu 0042, Fang Wan 0001, Jianbin Jiao, Xiangyang Ji, Qixiang Ye |
ICCV | 3 |
| 2019 | C-MIDN: Coupled Multiple Instance Detection Network With Segmentation Guidance for Weakly Supervised Object DetectionabstractWeakly supervised object detection (WSOD) that only needs image-level annotations has obtained much attention recently. By combining convolutional neural network with multiple instance learning method, Multiple Instance Detection Network (MIDN) has become the most popular method to address the WSOD problem and been adopted as the initial model in many works. We argue that MIDN inclines to converge to the most discriminative object parts, which limits the performance of methods based on it. In this paper, we propose a novel Coupled Multiple Instance Detection Network (C-MIDN) to address this problem. Specifically, we use a pair of MIDNs, which work in a complementary manner with proposal removal. The localization information of the MIDNs is further coupled to obtain tighter bounding boxes and localize multiple objects. We also introduce a Segmentation Guided Proposal Removal (SGPR) algorithm to guarantee the MIL constraint after the removal and ensure the robustness of C-MIDN. Through a simple implementation of the C-MIDN with online detector refinement, we obtain 53.6% and 50.3% mAP on the challenging PASCAL VOC 2007 and 2012 benchmarks respectively, which significantly outperform the previous state-of-the-arts. Gao Yan, Boxiao Liu, Nan Guo 0003, Xiaochun Ye, Fang Wan 0001, Haihang You, Dongrui Fan |
ICCV | 5 |
| 2019 | FreeAnchor: Learning to Match Anchors for Visual Object DetectionabstractModern CNN-based object detectors assign anchors for ground-truth objects under the restriction of object-anchor Intersection-over-Unit (IoU). In this study, we propose a learning-to-match approach to break IoU restriction, allowing objects to match anchors in a flexible manner. Our approach, referred to as FreeAnchor, updates hand-crafted anchor assignment to "free" anchor matching by formulating detector training as a maximum likelihood estimation (MLE) procedure. FreeAnchor targets at learning features which best explain a class of objects in terms of both classification and localization. FreeAnchor is implemented by optimizing detection customized likelihood and can be fused with CNN-based detectors in a plug-and-play manner. Experiments on MS-COCO demonstrate that FreeAnchor consistently outperforms the counterparts with significant margins. Xiaosong Zhang 0004, Fang Wan 0001, Chang Liu 0042, Rongrong Ji, Qixiang Ye |
NeurIPS | 2 |
| 2019 | Min-Entropy Latent Model for Weakly Supervised Object DetectionabstractWeakly supervised object detection is a challenging task when provided with image category supervision but required to learn, at the same time, object locations and object detectors. The inconsistency between the weak supervision and learning objectives introduces significant randomness to object locations and ambiguity to detectors. In this paper, a min-entropy latent model (MELM) is proposed for weakly supervised object detection. Min-entropy serves as a model to learn object locations and a metric to measure the randomness of object localization during learning. It aims to principally reduce the variance of learned instances and alleviate the ambiguity of detectors. MELM is decomposed into three components including proposal clique partition, object clique discovery, and object localization. MELM is optimized with a recurrent learning algorithm, which leverages continuation optimization to solve the challenging non-convexity problem. Experiments demonstrate that MELM significantly improves the performance of weakly supervised object detection, weakly supervised object localization, and image classification, against the state-of-the-art approaches. Fang Wan 0001, Pengxu Wei, Zhenjun Han, Jianbin Jiao, Qixiang Ye |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2018 | Min-Entropy Latent Model for Weakly Supervised Object DetectionabstractWeakly supervised object detection is a challenging task when provided with image category supervision but required to learn, at the same time, object locations and object detectors. The inconsistency between the weak supervision and learning objectives introduces randomness to object locations and ambiguity to detectors. In this paper, a min-entropy latent model (MELM) is proposed for weakly supervised object detection. Min-entropy is used as a metric to measure the randomness of object localization during learning, as well as serving as a model to learn object locations. It aims to principally reduce the variance of positive instances and alleviate the ambiguity of detectors. MELM is deployed as two sub-models, which respectively discovers and localizes objects by minimizing the global and local entropy. MELM is unified with feature learning and optimized with a recurrent learning algorithm, which progressively transfers the weak supervision to object locations. Experiments demonstrate that MELM significantly improves the performance of weakly supervised detection, weakly supervised localization, and image classification, against the state-of-the-art approaches. Fang Wan 0001, Pengxu Wei, Jianbin Jiao, Zhenjun Han, Qixiang Ye |
CVPR | 1 |
| 2017 | Correlated Topic Vector for Scene ClassificationabstractScene images usually involve semantic correlations, particularly when considering large-scale image data sets. This paper proposes a novel generative image representation, correlated topic vector, to model such semantic correlations. Oriented from the correlated topic model, correlated topic vector intends to naturally utilize the correlations among topics, which are seldom considered in the conventional feature encoding, e.g., Fisher vector, but do exist in scene images. It is expected that the involvement of correlations can increase the discriminative capability of the learned generative model and consequently improve the recognition accuracy. Incorporated with the Fisher kernel method, correlated topic vector inherits the advantages of Fisher vector. The contributions to the topics of visual words have been further employed by incorporating the Fisher kernel framework to indicate the differences among scenes. Combined with the deep convolutional neural network (CNN) features and Gibbs sampling solution, correlated topic vector shows great potential when processing large-scale and complex scene image data sets. Experiments on two scene image data sets demonstrate that correlated topic vector improves significantly the deep CNN features, and outperforms existing Fisher kernel-based features. Pengxu Wei, Fang Wan 0001, Yi Zhu 0004, Jianbin Jiao, Qixiang Ye |
IEEE Trans. Image Process. | 3 |
| 2016 | Weakly supervised object detection with correlation and part suppressionabstractIn weakly supervised object detection, conventional methods treat object location in each image as a latent variable and use non-convex optimization to solve the latent variable. However, as the optimization objective is image-level instead of sample-level, the learning procedure tends to choose object parts as false positive samples. Furthermore, when multiple classes of objects appear in the same images, the models could invite class-correlations and lose discriminative capability. In this paper, we propose a simple but effective suppression strategy that mines hard negative samples in the learning procedure to ease the above problems. We propose using a spatial-voting strategy to help finding negative samples to suppress the impact of object parts. We also use regions from class-correlated images as negative samples to suppress the impact of class-correlations. Experiments show that our approach significantly improves the baseline by 6% and achieves state-of-the-art performance. Fang Wan 0001, Pengxu Wei, Zhenjun Han, Kun Fu 0001, Qixiang Ye |
ICIP | 1 |
| 2016 | Collective motion pattern inference via Locally Consistent Latent Dirichlet Allocation
Jialing Zou, Qixiang Ye, Yanting Cui, Fang Wan 0001, Kun Fu 0001, Jianbin Jiao |
Neurocomputing | 4 |
| 2014 | A cluster specific latent dirichlet allocation model for trajectory clustering in crowded videosabstractTrajectory analysis in crowded video scenes is challenging as trajectories obtained by existing tracking algorithms are often fragmented. In this paper, we propose a new approach to do trajectory inference and clustering on fragmented trajectories, by exploring a cluster specific Latent Dirichlet Allocation(CLDA) model. LDA models are widely used to learn middle level trajectory features and perform trajectory inference. However, they often require scene priors in the learning or inference process. Our cluster specific LDA model addresses this issue by using manifold based clustering as initialization and iterative statistical inference as optimization. The output middle level features of CLDA are input to a clustering algorithm to obtain trajectory clusters. Experiments on a public dataset show the effectiveness of our approach. Jialing Zou, Yanting Cui, Fang Wan 0001, Qixiang Ye, Jianbin Jiao |
ICIP | 3 |