EDBT 2026 Demo / reviewers in the wild / expert
Qixiang Ye
dblp:06/4335
· DBLP profile ↗
192ranked-venue papers
16as first author
101since 2021 · last 2026
0000-0003-1215-6259ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 126 · 5 first-author · 73 since 2021Graphics, computer vision, multimedia, augmented reality and games · 119 · 10 first-author · 54 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 3 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Discrepancy-Controlled Region-Adaptive Learning: Handling Intra-Domain Bias for Crowd CountingabstractCrowd counting in congested scenarios remains challenging, when required to handle“intra-domain bias”—the significant variation in crowd density across regions within each image. In this study, we propose a novel Discrepancy-Controlled Region-Adaptive Learning (DC-RA) method which leverages a divide-and-conquer strategy, transforming the complex problem of image-level crowd counting into a series of more manageable regional tasks. Specifically, we propose a Discrepancy-Controlled Adaptive Partition (DCAP) module, to divide each image to regions that adapt to the varying density levels controlled by discrepancy of crowd density. To specify features for each region, the Region-wise Adaptive Learning (RAL) module is then introduced by incorporating the Mixture-of-Experts (MoE) framework, which involves using a routing module to select the most suitable expert for each region. This dynamic selection process ensures that each region benefits from tailored optimization based on its specific characteristics, leading to more precise density estimates. To ensure that each expert captures the distinct characteristics of various regions, we further incorporate a region-level counting loss for optimization. Experiments show that DC-RA reduces the Mean Absolute Errors (MAE) by 2.5 and 4.1 compared with the state-of-the-art method on JHU-CROWD++ and NWPU, respectively, significantly enhancing the model’s robustness and accuracy across varying crowd densities. Mingyue Guo 0001, Zimo Liu, Yaowei Wang 0001, Qixiang Ye |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | EinsPT: Efficient Instance-Aware Pre-Training of Vision Foundation ModelsabstractIn this study, we introduce EinsPT, an efficient instance-aware pre-training paradigm designed to reduce the transfer gap between vision foundation models and downstream instance-level tasks. Unlike conventional image-level pre-training that relies solely on unlabeled images, EinsPT leverages both image reconstruction and instance annotations to learn representations that are spatially coherent and instance discriminative. To achieve this efficiently, we propose a proxy-foundation architecture that decouples high-resolution and low-resolution learning: the foundation model processes masked low-resolution images for global semantics, while a lightweight proxy model operates on complete high-resolution images to preserve fine-grained details. The two branches are jointly optimized through reconstruction and instance-level prediction losses on fused features. Extensive experiments demonstrate that EinsPT consistently enhances recognition accuracy across various downstream tasks with substantially reduced computational cost, while qualitative results further reveal improved instance perception and completeness in visual representations. Code is available at github.com/feufhd/EinsPT. Zhaozhi Wang, Yunjie Tian, Lingxi Xie, Yaowei Wang 0001, Qixiang Ye |
IEEE Trans. Image Process. | 5 |
| 2026 | Expandable Residual Approximation for Knowledge DistillationabstractKnowledge distillation (KD) aims to transfer knowledge from a large-scale teacher model to a lightweight one, significantly reducing computational and storage requirements. However, the inherent learning capacity gap between the teacher and student often hinders the sufficient transfer of knowledge, motivating numerous studies to address this challenge. Inspired by the progressive approximation principle in the Stone-Weierstrass theorem, we propose expandable residual approximation (ERA), a novel KD method that decomposes the approximation of residual knowledge into multiple steps, reducing the difficulty of mimicking the teacher's representation through a divide-and-conquer approach. Specifically, ERA employs a multibranched residual network (MBRNet) to implement this residual knowledge decomposition. Additionally, a teacher weight integration (TWI) strategy is introduced to mitigate the capacity disparity by reusing the teacher's head weights. Extensive experiments show that ERA improves the Top-1 accuracy on ImageNet classification benchmark by 1.41% and the AP on the MS COCO object detection benchmark by 1.40, as well as achieving leading performance across computer vision tasks. Zhaoyi Yan, Binghui Chen, Yunfan Liu 0001, Qixiang Ye |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2025 | ChatterBox: Multimodal Referring and Grounding with Chain-of-QuestionsabstractIn this study, we establish a benchmark and a baseline approach for Multimodal referring and grounding with Chain-of-Questions (MCQ), opening up a promising direction for ‘logical’ multimodal dialogues. The newly collected dataset, named CB-300K, spans challenges including probing dialogues with spatial relationship among multiple objects, consistent reasoning, and complex question chains. The baseline approach, termed ChatterBox, involves a modularized design and a referent feedback mechanism to ensure logical coherence in continuous referring and grounding tasks. This design reduces the risk of referential confusion, simplifies the training process, and presents validity in retaining the language model’s generation ability. Experiments show that ChatterBox demonstrates superiority in MCQ both quantitatively and qualitatively, paving a new path towards multimodal dialogue scenarios with logical interactions. Yunjie Tian, Tianren Ma, Lingxi Xie, Qixiang Ye |
AAAI | 4 |
| 2025 | Timestep Embedding Tells: It's Time to Cache for Video Diffusion ModelabstractAs a fundamental backbone for video generation, diffusion models are challenged by low inference speed due to the sequential nature of denoising. Previous methods speed up the models by caching and reusing model outputs at uniformly selected timesteps. However, such a strategy neglects the fact that differences among model outputs are not uniform across timesteps, which hinders selecting the appropriate model outputs to cache, leading to a poor balance between inference efficiency and visual quality. In this study, we introduce Timestep Embedding Aware Cache (TeaCache), a training-free caching approach that estimates and leverages the fluctuating differences among model outputs across timesteps. Rather than directly using the time-consuming model outputs, TeaCache focuses on model inputs, which have a strong correlation with the modeloutputs while incurring negligible computational cost. TeaCache first modulates the noisy inputs using the timestep embeddings to ensure their differences better approximating those of model outputs. TeaCache then introduces a rescaling strategy to refine the estimated differences and utilizes them to indicate output caching. Experiments show that TeaCache achieves up to 4.41× acceleration over Open-Sora-Plan with negligible (-0.07% Vbench score) degradation of visual quality. Feng Liu 0050, Shiwei Zhang 0001, Yujie Wei 0001, Haonan Qiu, Yuzhong Zhao, Yingya Zhang, Qixiang Ye, Fang Wan 0001 |
CVPR | 8 |
| 2025 | Adaptive Keyframe Sampling for Long Video UnderstandingabstractMultimodal large language models (MLLMs) have enabled open-world visual understanding by injecting visual input as extra tokens into large language models (LLMs) as contexts. However, when the visual input changes from a single image to a long video, the above paradigm encounters difficulty because the vast amount of video tokens has significantly exceeded the maximal capacity of MLLMs. Therefore, existing video-based MLLMs are mostly established upon sampling a small portion of tokens from input data, which can cause key information to be lost and thus produce incorrect answers. This paper presents a simple yet effective algorithm named Adaptive Keyframe Sampling (AKS). It inserts a plug-and-play module known as keyframe selection, which aims to maximize the useful information with a fixed number of video tokens. We formulate keyframe selection as an optimization involving (1) the relevance between the keyframes and the prompt, and (2) the coverage of the keyframes over the video, and present an adaptive algorithm to approximate the best solution. Experiments on two long video understanding benchmarks validate that AKS improves video QA accuracy (beyond strong baselines) upon selecting informative keyframes. Our study reveals the importance of information pre-filtering in video-based MLLMs. Our codes are available at https://github.com/ncTimTang/AKS Jihao Qiu, Lingxi Xie, Yunjie Tian, Jianbin Jiao, Qixiang Ye |
CVPR | 6 |
| 2025 | Building Vision Models upon Heat ConductionabstractVisual representation models leveraging attention mechanisms are challenged by significant computational overhead, particularly when pursuing large receptive fields. In this study, we aim to mitigate this challenge by introducing the Heat Conduction Operator (HCO) built upon the physical heat conduction principle. HCO conceptualizes image patches as heat sources and models their correlations through adaptive thermal energy diffusion, enabling robust visual representations. HCO enjoys a computational complexity of O(N1.5), as it can be implemented using discrete cosine transformation (DCT) operations. HCO is plug-and-play, combining with deep learning backbones produces visual representation models (termed vHeat) with global receptive fields. Experiments across vision tasks demonstrate that, beyond the stronger performance, vHeat achieves up to a 3× throughput, 80% less GPU memory allocation, and 35% fewer computational FLOPs compared to the Swin-Transformer. Code is available at https://github.com/MzeroMiko/vHeat and https://openi.pcl.ac.cn/georgew/vHeat. Zhaozhi Wang, Yunjie Tian, Yunfan Liu 0001, Yaowei Wang 0001, Qixiang Ye |
CVPR | 6 |
| 2025 | DynRefer: Delving into Region-level Multimodal Tasks via Dynamic ResolutionabstractOne fundamental task of multimodal models is to translate referred image regions to human preferred language descriptions. Existing methods, however, ignore the resolution adaptability needs of different tasks, which hinders them to find out precise language descriptions. In this study, we propose a DynRefer approach, to pursue high-accuracy region-level referring through mimicking the resolution adaptability of human visual cognition. During training, DynRefer stochastically aligns language descriptions of multimodal tasks with images of multiple resolutions, which are constructed by nesting a set of random views around the referred region. During inference, DynRefer performs selectively multimodal referring by sampling proper region representations for tasks from the nested views based on image and task priors. This allows the visual information for referring to better match human preferences, thereby improving the representational adaptability of region-level multimodal models. Experiments show that DynRefer brings mutual improvement upon broad tasks including region-level captioning, open-vocabulary region recognition and attribute detection. Furthermore, DynRefer achieves state-of-the-art results on multiple region-level multimodal tasks using a single model. Code is available at https://github.com/callsys/DynRefer. Yuzhong Zhao, Feng Liu 0050, Mingxiang Liao, Chen Gong 0005, Qixiang Ye, Fang Wan 0001 |
CVPR | 6 |
| 2025 | RS-vHeat: Heat Conduction Guided Efficient Remote Sensing Foundation ModelabstractRemote sensing foundation models largely break away from the traditional paradigm of designing task-specific models, offering greater scalability across multiple tasks. However, they face challenges such as low computational efficiency and limited interpretability, especially when dealing with large-scale remote sensing images. To overcome these, we draw inspiration from heat conduction, a physical process modeling local heat diffusion. Building on this idea, we are the first to explore the potential of using the parallel computing model of heat conduction to simulate the local region correlations in high-resolution remote sensing images, and introduce RS-vHeat, an efficient multi-modal remote sensing foundation model. Specifically, RS-vHeat 1) applies the Heat Conduction Operator (HCO) with a complexity of $O(N^{1.5})$ and a global receptive field, reducing computational overhead while capturing remote sensing object structure information to guide heat diffusion; 2) learns the frequency distribution representations of various scenes through a self-supervised strategy based on frequency domain hierarchical masking and multi-domain reconstruction; 3) significantly improves efficiency and performance over state-of-the-art techniques across 4 tasks and 10 datasets. Compared to attention-based remote sensing foundation models, we reduce memory usage by 84\%, FLOPs by 24\% and improves throughput by 2.7 times. The code will be made publicly available. Huiyang Hu, Peijin Wang, Hanbo Bi, Boyuan Tong, Zhaozhi Wang, Wenhui Diao, Yingchao Feng, Ziqi Zhang 0010, Yaowei Wang 0001, Qixiang Ye, Kun Fu 0001, Xian Sun 0001 |
ICCV | 11 |
| 2025 | ClawMachine: Learning to Fetch Visual Tokens for Referential ComprehensionabstractAligning vision and language concepts at a finer level remains an essential topic of multimodal large language models (MLLMs), particularly for tasks such as referring and grounding. Existing methods, such as *proxy encoding* and *geometry encoding* genres, incorporate additional syntax to encode spatial information, imposing extra burdens when communicating between language with vision modules. In this study, we propose ClawMachine, offering a new methodology that explicitly notates each entity using **token collectives**—groups of visual tokens that collaboratively represent higher-level semantics. A hybrid perception mechanism is also explored to perceive and understand scenes from both discrete and continuous spaces. Our method unifies the prompt and answer of visual referential tasks without using additional syntax. By leveraging a joint vision-language vocabulary, ClawMachine integrates referring and grounding in an auto-regressive manner, demonstrating great potential with scaled up pre-training data. Experiments show that ClawMachine achieves superior performance on scene-level and referential understanding tasks with higher efficiency. It also exhibits the potential to integrate multi-source information for complex visual reasoning, which is beyond the capability of many MLLMs. Our code is available at https://github.com/martian422/ClawMachine. Tianren Ma, Lingxi Xie, Yunjie Tian, Boyu Yang 0002, Qixiang Ye |
ICLR | 5 |
| 2025 | YOLOv12: Attention-Centric Real-Time Object DetectorsabstractEnhancing the network architecture of the YOLO framework has been crucial for a long time. Still, it has focused on CNN-based improvements despite the proven superiority of attention mechanisms in modeling capabilities. This is because attention-based models cannot match the speed of CNN-based models. This paper proposes an attention-centric YOLO framework, namely YOLOv12, that matches the speed of previous CNN-based ones while harnessing the performance benefits of attention mechanisms. YOLOv12 surpasses popular real-time object detectors in accuracy with competitive speed. For example, YOLOv12-N achieves 40.5% mAP with an inference latency of 1.62 ms on a T4 GPU, outperforming advanced YOLOv10-N / YOLO11-N by 2.0%/1.1% mAP with a comparable speed. This advantage extends to other model scales. YOLOv12 also surpasses end-to-end real-time detectors that improve DETR, such as RT-DETRv2 / RT-DETRv3: YOLOv12-X beats RT-DETRv2-R101 / RT-DETRv3-R101 while running faster with fewer computations and parameters. See more comparisons in Figure 1. Source code is available at https://github.com/sunsmarterjie/yolov12. Yunjie Tian, Qixiang Ye, David S. Doermann |
NeurIPS | 2 |
| 2025 | Discriminatively Matched Part Tokens for Pointly Supervised Instance Segmentation
Zonghao Guo, Fang Wan 0001, Mingxiang Liao, Qixiang Ye |
Int. J. Comput. Vis. | 5 |
| 2025 | ClickTrack: Towards real-time interactive single object tracking
Kuiran Wang, Xuehui Yu, Wenwen Yu, Guorong Li, Xiangyuan Lan, Qixiang Ye, Jianbin Jiao, Zhenjun Han |
Pattern Recognit. | 6 |
| 2025 | Depth-Guided Texture Diffusion for Image Semantic SegmentationabstractDepth information provides valuable insights into the 3D structure, especially the outline of objects, which can be utilized to enhance semantic segmentation. However, a naive fusion of depth information can disrupt features and compromise accuracy due to the gap between depth and RGB modalities. In this work, we introduce a depth-guided texture diffusion approach that effectively tackles the outlined challenge. Our method extracts low-level features from edges and textures to create a texture image. This image is then selectively diffused across the depth map, enhancing structural information vital for precisely extracting object outlines. By integrating this enhanced depth map with the original RGB image into a joint feature embedding, our method effectively bridges the disparity between depth and RGB modalities, enabling more accurate semantic segmentation. We conduct comprehensive experiments on diverse, widely-used datasets covering various semantic segmentation tasks, including Camouflaged Object Detection (COD), Salient Object Detection (SOD), and indoor semantic segmentation. With source-free estimated depth or depth captured by depth cameras, our method consistently outperforms existing baselines and achieves new state-of-the-art results, demonstrating the effectiveness of our depth-guided texture diffusion for image semantic segmentation. The source code and datasets are publicly available athttps://github.com/Wistzz/Texture-Diffusion.git. Wei Sun 0056, Yuan Li 0055, Qixiang Ye, Jianbin Jiao, Yanzhao Zhou |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Language-Driven Visual Consensus for Zero-Shot Semantic SegmentationabstractThe pre-trained vision-language model, exemplified by CLIP, advances zero-shot semantic segmentation by aligning visual features with class embeddings through a transformer decoder to generate semantic masks. Despite its effectiveness, prevailing methods within this paradigm encounter challenges, including overfitting on seen classes and small fragmentation in segmentation masks. To mitigate these issues, we propose a Language-Driven Visual Consensus (LDVC) approach, fostering improved alignment of linguistic and visual information. Specifically, we leverage class embeddings as anchors due to their discrete and abstract nature, steering visual features toward class embeddings. Moreover, to achieve a more compact visual space, we introduce route attention into the transformer decoder to find visual consensus, thereby enhancing semantic consistency within the same object. Equipped with a vision-language prompting strategy, our approach significantly boosts the generalization capacity of segmentation models for unseen classes. Experimental results underscore the effectiveness of our approach, showcasing mIoU gains of 4.5% on the PASCAL VOC 2012 and 3.6% on the COCO-Stuff 164K for unseen classes compared with the state-of-the-art methods. Wei Ke 0003, Yi Zhu 0004, Xiaodan Liang, Jianzhuang Liu, Qixiang Ye, Tong Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | RingMoGPT: A Unified Remote Sensing Foundation Model for Vision, Language, and Grounded TasksabstractRecently, multimodal large language models (MLLMs) have shown excellent reasoning capabilities in various fields. Most of the existing remote sensing (RS) MLLMs solve image-level text generation problems (e.g., image captioning), but ignore the core issues of object-level recognition, location, and multitemporal changes in the field of RS. In this article, we propose RingMoGPT, a multimodal foundation model that unifies vision, language, and localization. Based on the idea of domain adaption, RingMoGPT can complete training by fine-tuning only a few parameters. To make the model capable of object detection and change captioning, we further propose a location- and instruction-aware querying transformer (Q-Former) and a change detection module, respectively. To improve the performance of RingMoGPT, we carefully design the pretraining dataset and the instruction-tuning dataset. The pretraining dataset contains over a half million high-quality image and text pairs, which are generated through a low-cost and efficient data generation paradigm. The instruction-tuning dataset contains more than 1.6 million question-answer pairs, including six downstream tasks: scene classification, object detection, visual question answering (VQA), image captioning, grounded image captioning, and change captioning. Our experiments show that RingMoGPT performs well on six tasks, especially its ability to analyze multitemporal data changes and identify dense objects. We also verified the model under a zero-shot setting, and the results show that the proposed RingMoGPT also has good generalization ability in the face of new data. Peijin Wang, Huiyang Hu, Boyuan Tong, Ziqi Zhang 0010, Fanglong Yao, Yingchao Feng, Zining Zhu 0004, Wenhui Diao, Qixiang Ye, Xian Sun 0001 |
IEEE Trans. Geosci. Remote. Sens. | 10 |
| 2025 | CC-Diff++: Spatially Controllable Text-to-Image Synthesis for Remote Sensing With Enhanced Contextual CoherenceabstractGenerating visually realistic remote sensing (RS) images requires maintaining semantic coherence between objects and their surrounding environments. However, existing image synthesis methods prioritize foreground controllability while oversimplifying backgrounds into plain or generic textures. This oversight neglects the crucial interaction between foreground and background elements, resulting in semantic inconsistencies in RS scenarios. To address this challenge, we propose CC-Diff++, a Diffusion Model-based approach for spatially controllable RS image synthesis with enhanced Context Coherence. To capture spatial interdependence, we propose a novel module named Co-Resampler, which employs an advanced masked attention mechanism to jointly extract features from both the foreground and background while modeling their mutual relationships. Furthermore, we introduce a text-to-layout prediction module powered by Large Language Models (LLMs) and a reference image retrieval mechanism for providing rich textural guidance, which work together to enable CC-Diff++ to generate outputs that are both more diverse and more realistic. Extensive experiments demonstrate that CC-Diff++ outperforms state-of-the-art methods in visual fidelity, semantic accuracy, and positional precision on multiple RS datasets. CC-Diff++ also shows strong trainability, improving detection accuracy by 2.04 mAP on DOTA and 11.81 mAP on the HRSC dataset. Mu Zhang 0019, Yunfan Liu 0001, Yuzhong Zhao, Qixiang Ye |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | Intra- and Inter-Head Orthogonal Attention for Image CaptioningabstractMulti-head attention (MA), which allows the model to jointly attend to crucial information from diverse representation subspaces through its heads, has yielded remarkable achievement in image captioning. However, there is no explicit mechanism to ensure MA attends to appropriate positions in diverse subspaces, resulting in overfocused attention for each head and redundancy between heads. In this paper, we propose a novel Intra- and Inter-Head Orthogonal Attention (I2OA) to efficiently improve MA in image captioning by introducing a concise orthogonal regularization to heads. Specifically, Intra-Head Orthogonal Attention enhances the attention learning of MA by introducing orthogonal constraint to each head, which decentralizes the object-centric attention to more comprehensive content-aware attention. Inter-Head Orthogonal Attention reduces the heads redundancy by applying orthogonal constraint between heads, which enlarges the diversity of representation subspaces and improves the representation ability for MA. Moreover, the proposed I2OA is flexible to combine with various multi-head attention based image captioning methods and improve the performances without increasing model complexity and parameters. Experiments on the MS COCO dataset demonstrate the effectiveness of the proposed model. Xiaodan Zhang 0003, Aozhe Jia, Junzhong Ji, Liangqiong Qu, Qixiang Ye |
IEEE Trans. Image Process. | 5 |
| 2025 | Virtual Classification: Modulating Domain-Specific Knowledge for Multidomain Crowd CountingabstractMultidomain crowd counting aims to learn a general model for multiple diverse datasets. However, deep networks prefer modeling distributions of the dominant domains instead of all domains, which is known as domain bias. In this study, we propose a simple-yet-effective modulating domain-specific knowledge network (MDKNet) to handle the domain bias issue in multidomain crowd counting. MDKNet is achieved by employing the idea of "modulating," enabling deep network balancing and modeling different distributions of diverse datasets with little bias. Specifically, we propose an instance-specific batch normalization (IsBN) module, which serves as a base modulator to refine the information flow to be adaptive to domain distributions. To precisely modulating the domain-specific information, the domain-guided virtual classifier (DVC) is then introduced to learn a domain-separable latent space. This space is employed as an input guidance for the IsBN modulator, such that the mixture distributions of multiple datasets can be well treated. Extensive experiments performed on popular benchmarks, including Shanghai-tech A/B, QNRF, and NWPU validate the superiority of MDKNet in tackling multidomain crowd counting and the effectiveness for multidomain learning. Code is available at https://github.com/csguomy/MDKNet. Mingyue Guo 0001, Binghui Chen, Zhaoyi Yan, Yaowei Wang 0001, Qixiang Ye |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2025 | Proposal Distribution Calibration for Few-Shot Object DetectionabstractAdapting object detectors learned with sufficient supervision to novel classes under low data regimes is charming yet challenging. In few-shot object detection (FSOD), the two-step training paradigm is widely adopted to mitigate the severe sample imbalance, i.e., holistic pre-training on base classes, then partial fine-tuning in a balanced setting with all classes. Since unlabeled instances are suppressed as backgrounds in the base training phase, the learned region proposal network (RPN) is prone to produce biased proposals for novel instances, resulting in dramatic performance degradation. Unfortunately, the extreme data scarcity aggravates the proposal distribution bias, hindering the region of interest (RoI) head from evolving toward novel classes. In this brief, we introduce a simple yet effective proposal distribution calibration (PDC) approach to neatly enhance the localization and classification abilities of the RoI head by recycling its localization ability endowed in base training and enriching high-quality positive samples for semantic fine-tuning. Specifically, we sample proposals based on the base proposal statistics to calibrate the distribution bias and impose additional localization and classification losses upon the sampled proposals for fast expanding the base detector to novel classes. Experiments on the commonly used Pascal VOC and MS COCO datasets with explicit state-of-the-art performances justify the efficacy of our PDC for FSOD. Code is available at github.com/Bohao-Lee/PDC. Chang Liu 0047, Xiaozhong Chen, Xiangyang Ji, Qixiang Ye |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2025 | Hierarchical AttentionShift for Pointly Supervised Instance SegmentationabstractPointly supervised instance segmentation (PSIS) remains a challenging task when appearance variances across object parts cause semantic inconsistency. In this article, we propose a hierarchical AttentionShift approach, to solve the semantic inconsistency issue through exploiting the hierarchical nature of semantics and the flexibility of key-point representation. The estimation of hierarchical attention is defined upon key-point sets. The representative key points are iteratively estimated spatially and in the feature space to capture the fine-grained semantics and cover the full object extent. Hierarchical AttentionShift is performed at instance, part, and fine-grained levels, optimizing object semantics while promoting the conventional self-attention activation to hierarchical activation with local refinement. Experiments on PASCAL VOC 2012 Aug and MS-COCO 2017 benchmarks show that hierarchical AttentionShift improves the state-of-the-art (SOTA) method by 10.4% and 7.0% upon mean average precision (mAP)50, respectively. When applying hierarchical AttentionShift to the segment anything model (SAM), 9.4% AP improvement on the COCO test-dev is achieved. Hierarchical AttentionShift provides a fresh insight to regularize the self-attention mechanism for fine-grained vision tasks. The code is available at github.com/MingXiangL/AttentionShift. Mingxiang Liao, Fang Wan 0001, Zonghao Guo, Qixiang Ye |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2025 | Explicit Margin Equilibrium for Few-Shot Object DetectionabstractUnder low data regimes, few-shot object detection (FSOD) transfers related knowledge from base classes with sufficient annotations to novel classes with limited samples in a two-step paradigm, including base training and balanced fine-tuning. In base training, the learned embedding space needs to be dispersed with large class margins to facilitate novel class accommodation and avoid feature aliasing while in balanced fine-tuning properly concentrating with small margins to represent novel classes precisely. Although obsession with the discrimination and representation dilemma has stimulated substantial progress, explorations for the equilibrium of class margins within the embedding space are still in full swing. In this study, we propose a class margin optimization scheme, termed explicit margin equilibrium (EME), by explicitly leveraging the quantified relationship between base and novel classes. EME first maximizes base-class margins to reserve adequate space to prepare for novel class adaptation. During fine-tuning, it quantifies the interclass semantic relationships by calculating the equilibrium coefficients based on the assumption that novel instances can be represented by linear combinations of base-class prototypes. EME finally reweights margin loss using equilibrium coefficients to adapt base knowledge for novel instance learning with the help of instance disturbance (ID) augmentation. As a plug-and-play module, EME can also be applied to few-shot classification. Consistent performance gains upon various baseline methods and benchmarks validate the generality and efficacy of EME. The code is available at github.com/Bohao-Lee/EME. Chang Liu 0047, Xiaozhong Chen, Qixiang Ye, Xiangyang Ji |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2025 | Exploring Complicated Search Spaces With Interleaving-Free SamplingabstractConventional neural architecture search (NAS) algorithms typically work on search spaces with short-distance node connections. We argue that such designs, though safe and stable, are obstacles to exploring more effective network architectures. In this brief, we explore the search algorithm upon a complicated search space with long-distance connections and show that existing weight-sharing search algorithms fail due to the existence of interleaved connections (ICs). Based on the observation, we present a simple-yet-effective algorithm, termed interleaving-free neural architecture search (IF-NAS). We further design a periodic sampling strategy to construct subnetworks during the search procedure, avoiding the ICs to emerge in any of them. In the proposed search space, IF-NAS outperforms both random sampling and previous weight-sharing search algorithms by significant margins. It can also be well-generalized to the microcell-based spaces. This study emphasizes the importance of macrostructure and we look forward to further efforts in this direction. The code is available at github.com/sunsmarterjie/IFNAS. Yunjie Tian, Lingxi Xie, Jiemin Fang, Jianbin Jiao, Qixiang Ye, Qi Tian 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2025 | Unsupervised Domain Adaptation on Person Reidentification Via Dual-Level Asymmetric Mutual LearningabstractUnsupervised domain adaptation (UDA) person reidentification (Re-ID) aims to identify pedestrian images within an unlabeled target domain with an auxiliary labeled source-domain dataset. Many existing works attempt to recover reliable identity information by considering multiple homogeneous networks. And take these generated labels to train the model in the target domain. However, these homogeneous networks identify people in approximate subspaces and equally exchange their knowledge with others or their mean net to improve their ability, inevitably limiting the scope of available knowledge and putting them into the same mistake. This article proposes a dual-level asymmetric mutual learning (DAML) method to learn discriminative representations from a broader knowledge scope with diverse embedding spaces. Specifically, two heterogeneous networks mutually learn knowledge from asymmetric subspaces through the pseudo label generation in a hard distillation manner. The knowledge transfer between two networks is based on an asymmetric mutual learning (AML) manner. The teacher network learns to identify both the target and source domain while adapting to the target domain distribution based on the knowledge of the student. Meanwhile, the student network is trained on the target dataset and employs the ground-truth label through the knowledge of the teacher. Extensive experiments in Market-1501, CUHK-SYSU, and MSMT17 public datasets verified the superiority of DAML over state-of-the-arts (SOTA). Qiong Wu 0012, Jiahan Li, Pingyang Dai, Qixiang Ye, Liujuan Cao, Yongjian Wu 0001, Rongrong Ji |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Spatial Transform Decoupling for Oriented Object DetectionabstractVision Transformers (ViTs) have achieved remarkable success in computer vision tasks. However, their potential in rotation-sensitive scenarios has not been fully explored, and this limitation may be inherently attributed to the lack of spatial invariance in the data-forwarding process. In this study, we present a novel approach, termed Spatial Transform Decoupling (STD), providing a simple-yet-effective solution for oriented object detection with ViTs. Built upon stacked ViT blocks, STD utilizes separate network branches to predict the position, size, and angle of bounding boxes, effectively harnessing the spatial transform potential of ViTs in a divide-and-conquer fashion. Moreover, by aggregating cascaded activation masks (CAMs) computed upon the regressed parameters, STD gradually enhances features within regions of interest (RoIs), which complements the self-attention mechanism. Without bells and whistles, STD achieves state-of-the-art performance on the benchmark datasets including DOTA-v1.0 (82.24% mAP) and HRSC2016 (98.55% mAP), which demonstrates the effectiveness of the proposed method. Source code is available at https://github.com/yuhongtian17/Spatial-Transform-Decoupling. Hongtian Yu, Yunjie Tian, Qixiang Ye, Yunfan Liu 0001 |
AAAI | 3 |
| 2024 | Regressor-Segmenter Mutual Prompt Learning for Crowd CountingabstractCrowd counting has achieved significant progress by training regressors to predict instance positions. In heavily crowded scenarios, however, regressors are challenged by uncontrollable annotation variance, which causes density map bias and context information inaccuracy. In this study, we propose mutual prompt learning (mPrompt), which leverages a regressor and a segmenter as guidance for each other, solving bias and inaccuracy caused by annotation variance while distinguishing foreground from back-ground. In specific, mPrompt leverages point annotations to tune the segmenter and predict pseudo head masks in a way of point prompt learning. It then uses the predicted segmentation masks, which serve as spatial constraint, to rectify biased point annotations as context prompt learning. mPrompt defines a way of mutual information maximization from prompt learning, mitigating the impact of annotation variance while improving model accuracy. Experiments show that mPrompt significantly reduces the Mean Aver-age Error (MAE), demonstrating the potential to be general framework for down-stream vision tasks. Code is available at https://github.com/csguomy/mPrompt. Mingyue Guo 0001, Li Yuan 0007, Zhaoyi Yan, Binghui Chen, Yaowei Wang 0001, Qixiang Ye |
CVPR | 6 |
| 2024 | Ray Denoising: Depth-Aware Hard Negative Sampling for Multi-view 3D Object Detection
Feng Liu 0050, Tengteng Huang, Qianjing Zhang, Fang Wan 0001, Qixiang Ye, Yanzhao Zhou |
ECCV (49) | 7 |
| 2024 | ControlCap: Controllable Region-Level Captioning
Yuzhong Zhao, Zonghao Guo, Weijia Wu 0001, Chen Gong 0005, Qixiang Ye, Fang Wan 0001 |
ECCV (38) | 6 |
| 2024 | Grounding Multimodal Large Language Models to the WorldabstractWe introduce Kosmos-2, a Multimodal Large Language Model (MLLM), enabling new capabilities of perceiving object descriptions (e.g., bounding boxes) and grounding text to the visual world. Specifically, we represent text spans (i.e., referring expressions and noun phrases) as links in Markdown, i.e., [text span](bounding boxes), where object descriptions are sequences of location tokens. To train the model, we construct a large-scale dataset about grounded image-text pairs (GrIT) together with multimodal corpora. In addition to the existing capabilities of MLLMs (e.g., perceiving general modalities, following instructions, and performing in-context learning), Kosmos-2 integrates the grounding capability to downstream applications, while maintaining the conventional capabilities of MLLMs (e.g., perceiving general modalities, following instructions, and performing in-context learning). Kosmos-2 is evaluated on a wide range of tasks, including (i) multimodal grounding, such as referring expression comprehension and phrase grounding, (ii) multimodal referring, such as referring expression generation, (iii) perception-language tasks, and (iv) language understanding and generation. This study sheds a light on the big convergence of language, multimodal perception, and world modeling, which is a key step toward artificial general intelligence. Code can be found in [https://aka.ms/kosmos-2](https://aka.ms/kosmos-2). Zhiliang Peng, Wenhui Wang 0003, Li Dong 0004, Yaru Hao, Shaohan Huang, Shuming Ma, Qixiang Ye, Furu Wei |
ICLR | 7 |
| 2024 | Kepler codebookabstractA codebook designed for learning discrete distributions in latent space has demonstrated state-of-the-art results on generation tasks. This inspires us to explore what distribution of codebook is better. Following the spirit of Kepler's Conjecture, we cast the codebook training as solving the sphere packing problem and derive a Kepler codebook with a compact and structured distribution to obtain a codebook for image representations. Furthermore, we implement the Kepler codebook training by simply employing this derived distribution as regularization and using the codebook partition method. We conduct extensive experiments to evaluate our trained codebook for image reconstruction and generation on natural and human face datasets, respectively, achieving significant performance improvement. Besides, our Kepler codebook has demonstrated superior performance when evaluated across datasets and even for reconstructing images with different resolutions. Our trained models and source codes will be publicly released. Junrong Lian, Ziyue Dong, Pengxu Wei, Wei Ke 0003, Chang Liu 0030, Qixiang Ye, Xiangyang Ji, Liang Lin 0004 |
ICML | 6 |
| 2024 | Evaluation of Text-to-Video Generation Models: A Dynamics PerspectiveabstractComprehensive and constructive evaluation protocols play an important role when developing sophisticated text-to-video (T2V) generation models. Existing evaluation protocols primarily focus on temporal consistency and content continuity, yet largely ignore dynamics of video content. Such dynamics is an essential dimension measuring the visual vividness and the honesty of video content to text prompts. In this study, we propose an effective evaluation protocol, termed DEVIL, which centers on the dynamics dimension to evaluate T2V generation models, as well as improving existing evaluation metrics. In practice, we define a set of dynamics scores corresponding to multiple temporal granularities, and a new benchmark of text prompts under multiple dynamics grades. Upon the text prompt benchmark, we assess the generation capacity of T2V models, characterized by metrics of dynamics ranges and T2V alignment. Moreover, we analyze the relevance of existing metrics to dynamics metrics, improving them from the perspective of dynamics. Experiments show that DEVIL evaluation metrics enjoy up to about 90\% consistency with human ratings, demonstrating the potential to advance T2V generation models. Mingxiang Liao, Hannan Lu, Qixiang Ye, Wangmeng Zuo, Fang Wan 0001, Tianyu Wang 0028, Yuzhong Zhao, Jingdong Wang 0001, Xinyu Zhang 0017 |
NeurIPS | 3 |
| 2024 | VMamba: Visual State Space ModelabstractDesigning computationally efficient network architectures remains an ongoing necessity in computer vision. In this paper, we adapt Mamba, a state-space language model, into VMamba, a vision backbone with linear time complexity. At the core of VMamba is a stack of Visual State-Space (VSS) blocks with the 2D Selective Scan (SS2D) module. By traversing along four scanning routes, SS2D bridges the gap between the ordered nature of 1D selective scan and the non-sequential structure of 2D vision data, which facilitates the collection of contextual information from various sources and perspectives. Based on the VSS blocks, we develop a family of VMamba architectures and accelerate them through a succession of architectural and implementation enhancements. Extensive experiments demonstrate VMamba’s
promising performance across diverse visual perception tasks, highlighting its superior input scaling efficiency compared to existing benchmark models. Source code is available at https://github.com/MzeroMiko/VMamba Yunjie Tian, Yuzhong Zhao, Hongtian Yu, Lingxi Xie, Yaowei Wang 0001, Qixiang Ye, Jianbin Jiao, Yunfan Liu 0001 |
NeurIPS | 7 |
| 2024 | Artemis: Towards Referential Understanding in Complex VideosabstractVideos carry rich visual information including object description, action, interaction, etc., but the existing multimodal large language models (MLLMs) fell short in referential understanding scenarios such as video-based referring. In this paper, we present Artemis, an MLLM that pushes video-based referential understanding to a finer level. Given a video, Artemis receives a natural-language question with a bounding box in any video frame and describes the referred target in the entire video. The key to achieving this goal lies in extracting compact, target-specific video features, where we set a solid baseline by tracking and selecting spatiotemporal features from the video. We train Artemis on the newly established ViderRef45K dataset with 45K video-QA pairs and design a computationally efficient, three-stage training procedure. Results are promising both quantitatively and qualitatively. Additionally, we show that Artemis can be integrated with video grounding and text summarization tools to understand more complex scenarios. Code and data are available at https://github.com/NeurIPS24Artemis/Artemis. Jihao Qiu, Lingxi Xie, Tianren Ma, Pengyu Yan, David S. Doermann, Qixiang Ye, Yunjie Tian |
NeurIPS | 8 |
| 2024 | Overcomplete-to-sparse representation learning for few-shot class-incremental learning
Mengying Fu, Binghao Liu, Tianren Ma, Qixiang Ye |
Multim. Syst. | 4 |
| 2024 | Fast-iTPN: Integrally Pre-Trained Transformer Pyramid Network With Token MigrationabstractWe propose integrally pre-trained transformer pyramid network (iTPN), towards jointly optimizing the network backbone and the neck, so that transfer gap between representation models and downstream tasks is minimal. iTPN is born with two elaborated designs: 1) The first pre-trained feature pyramid upon vision transformer (ViT). 2) Multi-stage supervision to the feature pyramid using masked feature modeling (MFM). iTPN is updated to Fast-iTPN, reducing computational memory overhead and accelerating inference through two flexible designs. 1) Token migration: dropping redundant tokens of the backbone while replenishing them in the feature pyramid without attention operations. 2) Token gathering: reducing computation cost caused by global attention by introducing few gathering tokens. The base/large-level Fast-iTPN achieve 88.75%/89.5% top-1 accuracy on ImageNet-1 K. With 1× training schedule using DINO, the base/large-level Fast-iTPN achieves 58.4%/58.8% box AP on COCO object detection, and a 57.5%/58.7% mIoU on ADE20 K semantic segmentation using MaskDINO. Fast-iTPN can accelerate the inference procedure by up to 70%, with negligible performance loss, demonstrating the potential to be a powerful backbone for downstream vision tasks. Yunjie Tian, Lingxi Xie, Jihao Qiu, Jianbin Jiao, Yaowei Wang 0001, Qi Tian 0001, Qixiang Ye |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2024 | CPR++: Object Localization via Single Coarse Point SupervisionabstractPoint-based object localization (POL), which pursues high-performance object sensing under low-cost data annotation, has attracted increased attention. However, the point annotation mode inevitably introduces semantic variance due to the inconsistency of annotated points. Existing POL heavily rely on strict annotation rules, which are difficult to define and apply, to handle the problem. In this study, we propose coarse point refinement (CPR), which to our best knowledge is the first attempt to alleviate semantic variance from an algorithmic perspective. CPR reduces the semantic variance by selecting a semantic centre point in a neighbourhood region to replace the initial annotated point. Furthermore, We design a sampling region estimation module to dynamically compute a sampling region for each object and use a cascaded structure to achieve end-to-end optimization. We further integrate a variance regularization into the structure to concentrate the predicted scores, yielding CPR++. We observe that CPR++ can obtain scale information and further reduce the semantic variance in a global region, thus guaranteeing high-performance object localization. Extensive experiments on four challenging datasets validate the effectiveness of both CPR and CPR++. We hope our work can inspire more research on designing algorithms rather than annotation rules to address the semantic variance problem in POL. Xuehui Yu, Pengfei Chen 0004, Kuiran Wang, Xumeng Han, Guorong Li, Zhenjun Han, Qixiang Ye, Jianbin Jiao |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2024 | Self-supervised feature-gate coupling for dynamic network pruning
Chang Liu 0047, Jianbin Jiao, Qixiang Ye |
Pattern Recognit. | 4 |
| 2024 | Save the Tiny, Save the All: Hierarchical Activation Network for Tiny Object DetectionabstractTiny object detection (TOD) remains a challenging problem due to the extremely small size and weak feature presentations of tiny objects. Many effective methods have improved the detection of small objects below$32\times 32$pixels to some extent, but the performance is still poor for the tiny objects below$16\times 16$pixels. In this paper, we find that the aliasing between the features and object scales, namely feature-scale-aliasing, leads to the misalignment between feature subspaces and detection subspaces, and thus results in the interference of features, especially for tiny objects. To alleviate this, we propose a Hierarchical Activation (HA) method to obtain scale-specific feature subspaces by activating object features at different scales hierarchically. To this end, we design a Scale-Guided Feature Activation (SGFA) to decompose the original object-aliasing feature spaces into a group of scale-specific feature subspaces by scale-guided activation maps. Then, Scale-Specific Feature re-Coupling (SSFC) is used to enhance the feature subspaces by adaptively aggregating the feature subspaces from different groups. In addition, we propose to complement the scale-specific detailed information by a designed Detailed Information Compensation (DIC) method. Implementing HA, a multi-scale keypoint-based detector is constructed to improve the tiny object detection, referred to as Hierarchical Activation Network (HANet). Extensive experiments are carried out on three tiny object detection datasets, e.g., TinyPerson, AI-TOD, and TinyCOCO. Our HANet achieves 58.45%$AP_{50}^{all}$, 22.1%$AP$, and 15.76%$AP$on TinyPerson, AI-TOD, and TinyCOCO, respectively, showing a significant performance gain over the competitors. Guangqian Guo, Pengfei Chen 0004, Xuehui Yu, Zhenjun Han, Qixiang Ye, Shan Gao 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Generic-to-Specific Distillation of Masked AutoencodersabstractTo transfer the representation capacity of large pre-trained models to lightweight models, knowledge distillation has been widely explored. However, conventional single-stage distillation methods are prone to getting stuck in the transfer of task-specific knowledge, making it difficult to retain task-agnostic knowledge which is crucial for model generalization. In this study, we propose generic-to-specific distillation (G2SD), to boost lightweight models under the assistance of large models pre-trained by masked image modeling. In generic distillation, the decoder of a small model is encouraged to align feature predictions with that of a large model, so that task-agnostic knowledge can be transferred. In specific distillation, predictions of the small model are encouraged to be consistent with those of the large model, to guarantee task performance. G2SD is also applicable for heterogeneous settings(i.e., distilling from ViT to CNN). With G2SD, the ViT-Small model respectively achieves 98.9%, 98.4%, 99.3% and 98.9% accuracies when compared with its teachers (ViT-Base) for image classification, object detection, semantic segmentation and video recognition tasks. The lightweight ResNet models are improved to a new height on image classification task. The code is available at github.com/pengzhiliang/G2SD. Zhiliang Peng, Li Dong 0004, Furu Wei, Qixiang Ye, Jianbin Jiao |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Few-Shot Object Detection in Remote-Sensing Images via Label-Consistent Classifier and Gradual RegressionabstractWith the abomination of time-consuming or even impractical large-scale labeling, few-shot object detection (FSOD) based on natural scenes has attracted extensive attention. However, directly migrating FSOD methods designed for natural images to large-size remote sensing images (RSIs) still remains challenges. 1) Labels of novel instances within the base dataset are inconsistently assigned between the base training and the few-shot fine-tuning stage, which confuses the detector and leads to significant performance degradation over novel classes. 2) The region proposal network (RPN) of detectors cannot provide sufficient high-quality proposals for remote sensing objects with various aspect ratios and irregular shapes, leading to decreased detection performance. To tackle these issues, we specify a novel few-shot object detector for RSIs, to avoid the significant performance degradation caused by inconsistent labeling assignments, as well as efficiently leveraging the novel instances that existed in the base dataset. Furthermore, the proposed detector utilizes a coarse-to-fine regression method with an enhanced feature extractor called Gradual RPN to improve the recall of RPN. Experiments on a newly constructed few-shot detection benchmark show that our approach improves the mAP of novel classes by up to 8.4% and the average recall of RPN by up to 12.3%. The source code is available at here. Yanxing Liu, Zongxu Pan, Bingchen Zhang, Qixiang Ye |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2024 | Beyond Instance Discrimination: Relation-Aware Contrastive Self-Supervised LearningabstractContrastive self-supervised learning (CSL) based on instance discrimination typically attracts positive samples while repelling negatives to learn representations with pre-defined binary self-supervision. However, vanilla CSL is inadequate in modeling sophisticated instance relations, limiting the learned model to retain fine semantic structure. On the one hand, samples with the same semantic category are inevitably pushed away as negatives. On the other hand, differences among samples cannot be captured. In this paper, we present relation-aware contrastive self-supervised learning (ReCo) to integrate instance relations, i.e., global distribution relation and local interpolation relation, into the CSL framework in a plug-and-play fashion. Specifically, we align similarity distributions calculated between the positive anchor views and the negatives at the global level to exploit diverse similarity relations among instances. Local-level interpolation consistency between the pixel space and the feature space is applied to quantitatively model the feature differences of samples with distinct apparent similarities. Through explicitly instance relation modeling, our ReCo avoids irrationally pushing away semantically identical samples and carves a well-structured feature space. Extensive experiments conducted on commonly used benchmarks justify that our ReCo consistently gains remarkable performance improvements. Yifei Zhang 0005, Chang Liu 0047, Yu Zhou 0015, Weiping Wang 0005, Qixiang Ye, Xiangyang Ji |
IEEE Trans. Multim. | 5 |
| 2024 | TS-CAM: Token Semantic Coupled Attention Map for Weakly Supervised Object LocalizationabstractWeakly supervised object localization (WSOL), which trains object localization models using solely image category annotations, remains a challenging problem. Existing approaches based on convolutional neural networks (CNNs) tend to miss full object extent while activating discriminative object parts. Based on our analysis, this is caused by CNN's intrinsic characteristics, which experiences difficulty to capture object semantics at long distances. In this article, we introduce the vision transformer to WSOL, with the aim to capture long-range semantic dependency of features by leveraging transformer's cascaded self-attention mechanism. We propose the token semantic coupled attention map (TS-CAM) method, which first decomposes class-aware semantics and then couples the semantics with attention maps for semantic-aware activation. To capture object semantics at long distances and avoid partial activation, TS-CAM performs spatial embedding by partitioning an image to a set of patch tokens. To incorporate object category information to patch tokens, TS-CAM reallocates category-related semantics to each patch token. The patch tokens are finally coupled with attention maps which are semantic-agnostic to perform semantic-aware object localization. By introducing semantic tokens to produce semantic-aware attention maps, we further explore the capability of TS-CAM for multicategory object localization. Experiments show that TS-CAM outperforms its CNN-CAM counterpart by 11.6% and 28.9% on ILSVRC and CUB-200-2011 datasets, respectively, improving the state-of-the-art with large margins. TS-CAM also demonstrates superiority for multicategory object localization on the Pascal VOC dataset. The code is available at github.com/yuanyao366/ts-cam-extension. Fang Wan 0001, Wei Gao 0050, Xingjia Pan, Zhiliang Peng, Qi Tian 0001, Qixiang Ye |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2023 | Generic-to-Specific Distillation of Masked AutoencodersabstractLarge vision Transformers (ViTs) driven by self-supervised pre-training mechanisms achieved unprecedented progress. Lightweight ViT models limited by the model capacity, however, benefit little from those pre-training mechanisms. Knowledge distillation defines a paradigm to transfer representations from large (teacher) models to small (student) ones. However, the conventional single-stage distillation easily gets stuck on task-specific transfer, failing to retain the task-agnostic knowledge crucial for model generalization. In this study, we propose generic-to-specific distillation (G2SD), to tap the potential of small ViT models under the supervision of large models pretrained by masked autoencoders. In generic distillation, decoder of the small model is encouraged to align feature predictions with hidden representations of the large model, so that task-agnostic knowledge can be transferred. In specific distillation, predictions of the small model are constrained to be consistent with those of the large model, to transfer task-specific features which guarantee task performance. With G2SD, the vanilla ViT-Small model respectively achieves 98.7%, 98.1% and 99.3% the performance of its teacher (ViT-Base) for image classification, object detection, and semantic segmentation, setting a solid baseline for two-stage vision distillation. Code will be available at https://github.com/pengzhiliang/G2SD Zhiliang Peng, Li Dong 0004, Furu Wei, Jianbin Jiao, Qixiang Ye |
CVPR | 6 |
| 2023 | Integrally Pre-Trained Transformer Pyramid NetworksabstractIn this paper, we present an integral pre-training framework based on masked image modeling (MIM). We advocate for pre-training the backbone and neck jointly so that the transfer gap between MIM and downstream recognition tasks is minimal. We make two technical contributions. First, we unify the reconstruction and recognition necks by inserting a feature pyramid into the pre-training stage. Second, we complement mask image modeling (MIM) with masked feature modeling (MFM) that offers multi-stage supervision to the feature pyramid. The pre-trained models, termed integrally pre-trained transformer pyramid networks (iTPNs), serve as powerful foundation models for visual recognition. In particular, the base/large-level iTPN achieves an 86.2%/87.8% top-1 accuracy on ImageNet-1K, a 53.2%/55.6% box AP on COCO object detection with 1× training schedule using Mask-RCNN, and a 54.7%/57.7% mIoU on ADE20K semantic segmentation using UPerHead – all these results set new records. Our work inspires the community to work on unifying upstream pre-training and downstream fine-tuning tasks. Code is available at github.com/sunsmarterjie/iTPN. Yunjie Tian, Lingxi Xie, Zhaozhi Wang, Longhui Wei, Xiaopeng Zhang 0008, Jianbin Jiao, Yaowei Wang 0001, Qi Tian 0001, Qixiang Ye |
CVPR | 9 |
| 2023 | Multi-Agent Automated Machine LearningabstractIn this paper, we propose multi-agent automated machine learning (MA2ML) with the aim to effectively handle joint optimization of modules in automated machine learning (AutoML). MA2ML takes each machine learning module, such as data augmentation (AUG), neural architecture search (NAS), or hyper-parameters (HPO), as an agent and the final performance as the reward, to formulate a multi-agent reinforcement learning problem. MA2ML explicitly assigns credit to each agent according to its marginal contribution to enhance cooperation among modules, and incorporates off-policy learning to improve search efficiency. Theoretically, MA2ML guarantees monotonic improvement of joint optimization. Extensive experiments show that MA2ML yields the state-of-the-art top-1 accuracy on ImageNet under constraints of computational cost, e.g., 79.7%/80.5% with FLOPs fewer than 600M/800M. Exten\sive ablation studies verify the benefits of credit assignment and off-policy learning of MA2ML. Zhaozhi Wang, Kefan Su, Jian Zhang 0018, Huizhu Jia, Qixiang Ye, Zongqing Lu 0002 |
CVPR | 5 |
| 2023 | Integrally Migrating Pre-trained Transformer Encoder-decoders for Visual Object DetectionabstractModern object detectors have taken the advantages of backbone networks pre-trained on large scale datasets. Except for the backbone networks, however, other components such as the detector head and the feature pyramid network (FPN) remain trained from scratch, which hinders the generalization capacity of detectors. In this study, we propose to integrally migrate pre-trained transformer encoder-decoders (imTED) to a detector, constructing a feature extraction path which is "fully pre-trained" so that detectors’ generalization capacity is maximized. The essential differences between imTED with the baseline detector are twofold: (1) migrating the pre-trained transformer decoder to the detector head while removing the randomly initialized FPN from the feature extraction path; and (2) defining a multi-scale feature modulator (MFM) to enhance scale adaptability. Such designs not only reduce randomly initialized parameters significantly but also unify detector training with representation learning intendedly. Experiments on the MS COCO object detection dataset show that imTED consistently outperforms its counterparts by ~2.4 AP. Without bells and whistles, imTED improves the state-of-the-art of few-shot object detection by up to 7.6 AP. Code is released at https://github.com/LiewFeng/imTED. Feng Liu 0050, Xiaosong Zhang 0004, Zhiliang Peng, Zonghao Guo, Fang Wan 0001, Xiangyang Ji, Qixiang Ye |
ICCV | 7 |
| 2023 | Generative Prompt Model for Weakly Supervised Object LocalizationabstractWeakly supervised object localization (WSOL) remains challenging when learning object localization models from image category labels. Conventional methods that discriminatively train activation models ignore representative yet less discriminative object parts. In this study, we propose a generative prompt model (GenPromp), defining the first generative pipeline to localize less discriminative object parts by formulating WSOL as a conditional image denoising procedure. During training, GenPromp converts image category labels to learnable prompt embeddings which are fed to a generative model to conditionally recover the input image with noise and learn representative embeddings. During inference, GenPromp combines the representative embeddings with discriminative embeddings (queried from an off-the-shelf vision-language model) for both representative and discriminative capacity. The combined embeddings are finally used to generate multi-scale high-quality attention maps, which facilitate localizing full object extent. Experiments on CUB-200-2011 and ILSVRC show that GenPromp respectively outperforms the best discriminative models by 5.2% and 5.6% (Top-1 Loc), setting a solid baseline for WSOL with the generative model. Code is available at https://github.com/callsys/GenPromp. Yuzhong Zhao, Qixiang Ye, Weijia Wu 0001, Chunhua Shen, Fang Wan 0001 |
ICCV | 2 |
| 2023 | HiViT: A Simpler and More Efficient Design of Hierarchical Vision Transformer
Xiaosong Zhang 0004, Yunjie Tian, Lingxi Xie, Qi Dai 0001, Qixiang Ye, Qi Tian 0001 |
ICLR | 6 |
| 2023 | Towards Deviation-Robust Agent Navigation via Perturbation-Aware Contrastive LearningabstractVision-and-language navigation (VLN) asks an agent to follow a given language instruction to navigate through a real 3D environment. Despite significant advances, conventional VLN agents are trained typically under disturbance-free environments and may easily fail in real-world navigation scenarios, since they are unaware of how to deal with various possible disturbances, such as sudden obstacles or human interruptions, which widely exist and may usually cause an unexpected route deviation. In this paper, we present a model-agnostic training paradigm, called Progressive Perturbation-aware Contrastive Learning (PROPER) to enhance the generalization ability of existing VLN agents to the real world, by requiring them to learn towards deviation-robust navigation. Specifically, a simple yet effective path perturbation scheme is introduced to implement the route deviation, with which the agent is required to still navigate successfully following the original instruction. Since directly enforcing the agent to learn perturbed trajectories may lead to insufficient and inefficient training, a progressively perturbed trajectory augmentation strategy is designed, where the agent can self-adaptively learn to navigate under perturbation with the improvement of its navigation performance for each specific trajectory. For encouraging the agent to well capture the difference brought by perturbation and adapt to both perturbation-free and perturbation-based environments, a perturbation-aware contrastive learning mechanism is further developed by contrasting perturbation-free trajectory encodings and perturbation-based counterparts. Extensive experiments on the standard Room-to-Room (R2R) benchmark show that PROPER can benefit multiple state-of-the-art VLN baselines in perturbation-free scenarios. We further collect the perturbed path data to construct an introspection subset based on the R2R, called Path-Perturbed R2R (PP-R2R). The results on PP-R2R show unsatisfying robustness of popular VLN agents and the capability of PROPER in improving the navigation robustness under deviation. Bingqian Lin, Yanxin Long, Yi Zhu 0004, Fengda Zhu, Xiaodan Liang, Qixiang Ye, Liang Lin 0004 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2023 | Learnable Distribution Calibration for Few-Shot Class-Incremental LearningabstractFew-shot class-incremental learning (FSCIL) faces the challenges of memorizing old class distributions and estimating new class distributions given few training samples. In this study, we propose a learnable distribution calibration (LDC) approach, to systematically solve these two challenges using a unified framework. LDC is built upon a parameterized calibration unit (PCU), which initializes biased distributions for all classes based on classifier vectors (memory-free) and a single covariance matrix. The covariance matrix is shared by all classes, so that the memory costs are fixed. During base training, PCU is endowed with the ability to calibrate biased distributions by recurrently updating sampled features under supervision of real distributions. During incremental learning, PCU recovers distributions for old classes to avoid 'forgetting', as well as estimating distributions and augmenting samples for new classes to alleviate 'over-fitting' caused by the biased distributions of few-shot samples. LDC is theoretically plausible by formatting a variational inference procedure. It improves FSCIL's flexibility as the training procedure requires no class similarity priori. Experiments on CUB200, CIFAR100, and mini-ImageNet datasets show that LDC respectively outperforms the state-of-the-arts by 4.64%, 1.98%, and 3.97%. LDC's effectiveness is also validated on few-shot learning scenarios. Binghao Liu, Boyu Yang 0002, Lingxi Xie, Qi Tian 0001, Qixiang Ye |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2023 | Conformer: Local Features Coupling Global Representations for Recognition and DetectionabstractWith convolution operations, Convolutional Neural Networks (CNNs) are good at extracting local features but experience difficulty to capture global representations. With cascaded self-attention modules, vision transformers can capture long-distance feature dependencies but unfortunately deteriorate local feature details. In this paper, we propose a hybrid network structure, termed Conformer, to take both advantages of convolution operations and self-attention mechanisms for enhanced representation learning. Conformer roots in feature coupling of CNN local features and transformer global representations under different resolutions in an interactive fashion. Conformer adopts a dual structure so that local details and global dependencies are retained to the maximum extent. We also propose a Conformer-based detector (ConformerDet), which learns to predict and refine object proposals, by performing region-level feature coupling in an augmented cross-attention fashion. Experiments on ImageNet and MS COCO datasets validate Conformer's superiority for visual recognition and object detection, demonstrating its potential to be a general backbone network. Zhiliang Peng, Zonghao Guo, Yaowei Wang 0001, Lingxi Xie, Jianbin Jiao, Qi Tian 0001, Qixiang Ye |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2023 | Multiple Instance Differentiation Learning for Active Object DetectionabstractDespite the substantial progress of active learning for image recognition, there lacks a systematic investigation of instance-level active learning for object detection. In this paper, we propose to unify instance uncertainty calculation with image uncertainty estimation for informative image selection, creating a multiple instance differentiation learning (MIDL) method for instance-level active learning. MIDL consists of a classifier prediction differentiation module and a multiple instance differentiation module. The former leverages two adversarial instance classifiers trained on the labeled and unlabeled sets to estimate instance uncertainty of the unlabeled set. The latter treats unlabeled images as instance bags and re-estimates image-instance uncertainty using the instance classification model in a multiple instance learning fashion. Through weighting the instance uncertainty using instance class probability and instance objectness probability under the total probability formula, MIDL unifies the image uncertainty with instance uncertainty in the Bayesian theory framework. Extensive experiments validate that MIDL sets a solid baseline for instance-level active learning. On commonly used object detection datasets, it outperforms other state-of-the-art methods by significant margins, particularly when the labeled sets are small. Fang Wan 0001, Qixiang Ye, Tianning Yuan, Songcen Xu, Jianzhuang Liu, Xiangyang Ji, Qingming Huang |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Dynamic Support Network for Few-Shot Class Incremental LearningabstractFew-shot class-incremental learning (FSCIL) is challenged by catastrophically forgetting old classes and over-fitting new classes. Revealed by our analyses, the problems are caused by feature distribution crumbling, which leads to class confusion when continuously embedding few samples to a fixed feature space. In this study, we propose a Dynamic Support Network (DSN), which refers to an adaptively updating network with compressive node expansion to "support" the feature space. In each training session, DSN tentatively expands network nodes to enlarge feature representation capacity for incremental classes. It then dynamically compresses the expanded network by node self-activation to pursue compact feature representation, which alleviates over-fitting. Simultaneously, DSN selectively recalls old class distributions during incremental learning to support feature distributions and avoid confusion between classes. DSN with compressive node expansion and class distribution recalling provides a systematic solution for the problems of catastrophic forgetting and overfitting. Experiments on CUB, CIFAR-100, and miniImage datasets show that DSN significantly improves upon the baseline approach, achieving new state-of-the-arts. Boyu Yang 0002, Mingbao Lin, Binghao Liu, Xiaodan Liang, Rongrong Ji, Qixiang Ye |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2023 | CrossRectify: Leveraging disagreement for semi-supervised object detection
Chengcheng Ma, Xingjia Pan, Qixiang Ye, Fan Tang, Weiming Dong, Changsheng Xu |
Pattern Recognit. | 3 |
| 2023 | Rethinking Sampling Strategies for Unsupervised Person Re-IdentificationabstractUnsupervised person re-identification (re-ID) remains a challenging task. While extensive research has focused on the framework design and loss function, this paper shows that sampling strategy plays an equally important role. We analyze the reasons for the performance differences between various sampling strategies under the same framework and loss function. We suggest that deteriorated over-fitting is an important factor causing poor performance, and enhancing statistical stability can rectify this problem. Inspired by that, a simple yet effective approach is proposed, termed group sampling, which gathers samples from the same class into groups. The model is thereby trained using normalized group samples, which helps alleviate the negative impact of individual samples. Group sampling updates the pipeline of pseudo-label generation by guaranteeing that samples are more efficiently classified into the correct classes. It regulates the representation learning process, enhancing statistical stability for feature representation in a progressive fashion. Extensive experiments on Market-1501, DukeMTMC-reID and MSMT17 show that group sampling achieves performance comparable to state-of-the-art methods and outperforms the current techniques under purely camera-agnostic settings. Code has been available at https://github.com/ucas-vg/GroupSampling. Xumeng Han, Xuehui Yu, Guorong Li, Jian Zhao 0006, Gang Pan 0002, Qixiang Ye, Jianbin Jiao, Zhenjun Han |
IEEE Trans. Image Process. | 6 |
| 2023 | Transformer Sub-Patch Matching for High-Performance Visual Object TrackingabstractVisual tracking is a core component of intelligent transportation systems, especially for unmanned driving and road surveillance. Numerous convolutional neural network (CNN) trackers have achieved unprecedented performance. However, CNN features with regular spatial context relationships experience difficulty matching the rigid target templates when dramatic deformation and occlusion occur. In this paper, we propose a novel full Transformer Sub-patch Matching network for tracking (TSMtrack), which decomposes the tracked object into sub-patches, and interlaced matches the extracted sub-patches by leveraging the attention mechanism born with the Transformer. Roots in Transformer architecture, TSMtrack consists of image patch decomposition, sub-patch matching, and position prediction. Specifically, TSMtrack converts the whole frame into sub-patches and extracts the sub-patch features independently. By sub-patch matching and FFN-like prediction, TSMtrack enables independent similarity measurement between sub-patch features in an interlaced and iterative fashion. With a full Transformer pipeline implemented, we achieve a high-quality trade-off between tracking speed performance. Experiments on nine benchmarks demonstrate the effectiveness of our Transformer sub-patch matching framework. In particular, it realizes an AO of 75.6 on GOT-10K and SR of 57.9 on WebUAV-3M with 48 FPS on GPU RTX-2060s. Chuanming Tang, Qintao Hu, Gaofan Zhou, Jinzhen Yao, Jianlin Zhang 0001, Yongmei Huang, Qixiang Ye |
IEEE Trans. Intell. Transp. Syst. | 7 |
| 2023 | Anti-UAV: A Large-Scale Benchmark for Vision-Based UAV TrackingabstractUnmanned Aerial Vehicles (UAV) have many applications in both commerce and recreation. However, irresponsibly operated UAVs will pose a threat to public safety. Therefore, developing our understanding of UAVs and their uses is of particular interest. This paper considers tracking UAVs, which provide multifaceted information around location, paths and trajectories. To facilitate research on this topic, we introduce a new benchmark, herein referred to as Anti-UAV, which provides a novel direction for UAV tracking with more than 300 video pairs containing over 580 k manually annotated bounding boxes. Addressing anti-UAV research challenges could help to design anti-UAV systems, which in turn may improve surveillance. Accordingly, we have proposed a simple yet effective approach, called dual-flow semantic consistency (DFSC) is proposed for UAV tracking. Modulated by the semantic flow across video sequences, tracker learns more robust class-level semantic information and obtains more discriminative instance-level features. Experiments highlight significant performance gain with the proposed approach over state-of-the-art trackers and the challenging aspects of Anti-UAV. The Anti-UAV benchmark and the code for the proposed approach have been made publicly available athttps://github.com/ucas-vg/Anti-UAVandhttps://github.com/ZhaoJ9014/Anti-UAV. Kuiran Wang, Xiaoke Peng, Xuehui Yu, Qiang Wang 0051, Junliang Xing, Guorong Li, Guodong Guo, Qixiang Ye, Jianbin Jiao, Jian Zhao 0006, Zhenjun Han |
IEEE Trans. Multim. | 9 |
| 2023 | Self-Supervised Motion Perception for Spatiotemporal Representation LearningabstractIn this study, we propose a novel pretext task and a self-supervised motion perception (SMP) method for spatiotemporal representation learning. The pretext task is defined as video playback rate perception, which utilizes temporal dilated sampling to augment video clips to multiple duplicates of different temporal resolutions. The SMP method is built upon discriminative and generative motion perception models, which capture representations related to motion dynamics and appearance from video clips of multiple temporal resolutions in a collaborative fashion. To enhance the collaboration, we further propose difference and convolution motion attention (MA), which drives the generative model focusing on motion-related appearance, and leverage multiple granularity perception (MG) to extract accurate motion dynamics. Extensive experiments demonstrate SMP's effectiveness for video motion perception and state-of-the-art performance of self-supervised representation models upon target tasks, including action recognition and video retrieval. Code for SMP is available at github.com/yuanyao366/SMP. Chang Liu 0047, Dezhao Luo, Yu Zhou 0015, Qixiang Ye |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2022 | Object Localization under Single Coarse Point SupervisionabstractPoint-based object localization (POL), which pursues high-performance object sensing under low-cost data annotation, has attracted increased attention. However, the point annotation mode inevitably introduces semantic variance for the inconsistency of annotated points. Existing POL methods heavily reply on accurate keypoint annotations which are difficult to define. In this study, we propose a POL method using coarse point annotations, relaxing the supervision signals from accurate key points to freely spotted points. To this end, we propose a coarse point refinement (CPR) approach, which to our best knowledge is the first attempt to alleviate semantic variance from the perspective of algorithm. CPR constructs point bags, selects semantic-correlated points, and produces semantic center points through multiple instance learning (MIL). In this way, CPR defines a weakly supervised evolution procedure, which ensures training high-performance object localizer under coarse point supervision. Experimental results on COCO, DOTA and our proposed SeaPerson dataset validate the effectiveness of the CPR approach. The dataset and code will be available at https://github.com/ucas-vg/PointTinyBenchmark/ Xuehui Yu, Pengfei Chen 0004, Najmul Hassan, Guorong Li, Junchi Yan, Humphrey Shi, Qixiang Ye, Zhenjun Han |
CVPR | 8 |
| 2022 | Point-to-Box Network for Accurate Object Detection via Single Point Supervision
Pengfei Chen 0004, Xuehui Yu, Xumeng Han, Najmul Hassan, Kai Wang 0058, Jiachen Li 0003, Jian Zhao 0006, Humphrey Shi, Zhenjun Han, Qixiang Ye |
ECCV (9) | 10 |
| 2022 | End-to-End Weakly Supervised Object Detection with Sparse Proposal Evolution
Mingxiang Liao, Fang Wan 0001, Zhenjun Han, Jialing Zou, Yuze Wang 0004, Bailan Feng, Qixiang Ye |
ECCV (9) | 9 |
| 2022 | Adversarial Reinforced Instruction Attacker for Robust Vision-Language NavigationabstractLanguage instruction plays an essential role in the natural language grounded navigation tasks. However, navigators trained with limited human-annotated instructions may have difficulties in accurately capturing key information from the complicated instruction at different timesteps, leading to poor navigation performance. In this paper, we exploit to train a more robust navigator which is capable of dynamically extracting crucial factors from the long instruction, by using an adversarial attacking paradigm. Specifically, we propose a Dynamic Reinforced Instruction Attacker (DR-Attacker), which learns to mislead the navigator to move to the wrong target by destroying the most instructive information in instructions at different timesteps. By formulating the perturbation generation as a Markov Decision Process, DR-Attacker is optimized by the reinforcement learning algorithm to generate perturbed instructions sequentially during the navigation, according to a learnable attack score. Then, the perturbed instructions, which serve as hard samples, are used for improving the robustness of the navigator with an effective adversarial training strategy and an auxiliary self-supervised reasoning task. Experimental results on both Vision-and-Language Navigation (VLN) and Navigation from Dialog History (NDH) tasks show the superiority of our proposed method over state-of-the-art methods. Moreover, the visualization analysis shows the effectiveness of the proposed DR-Attacker, which can successfully attack crucial information in the instructions at different timesteps. Code is available at https://github.com/expectorlin/DR-Attacker. Bingqian Lin, Yi Zhu 0004, Yanxin Long, Xiaodan Liang, Qixiang Ye, Liang Lin 0004 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2022 | Learning to Match Anchors for Visual Object DetectionabstractModern CNN-based object detectors assign anchors for ground-truth objects under the restriction of object-anchor Intersection-over-Union (IoU). In this study, we propose a learning-to-match (LTM) method to break IoU restriction, allowing objects to match anchors in a flexible manner. LTM updates hand-crafted anchor assignment to "free" anchor matching by formulating detector training in the Maximum Likelihood Estimation (MLE) framework. During the training phase, LTM is implemented by converting the detection likelihood to anchor matching loss functions which are plug-and-play. Minimizing the matching loss functions drives learning and selecting features which best explain a class of objects with respect to both classification and localization. LTM is extended from anchor-based detectors to anchor-free detectors, validating the general applicability of learnable object-feature matching mechanism for visual object detection. Experiments on MS COCO dataset demonstrate that LTM detectors consistently outperform counterpart detectors with significant margins. The last but not the least, LTM requires negligible computational cost in both training and inference phases as it does not involve any additional architecture or parameter. Code has been made publicly available. Xiaosong Zhang 0004, Fang Wan 0001, Chang Liu 0047, Xiangyang Ji, Qixiang Ye |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2022 | Discrepant multiple instance learning for weakly supervised object detection
Wei Gao 0050, Fang Wan 0001, Jun Yue 0004, Songcen Xu, Qixiang Ye |
Pattern Recognit. | 5 |
| 2022 | Multi-View correlation distillation for incremental object detection
Dongbao Yang, Yu Zhou 0015, Aoting Zhang, Xurui Sun, Dayan Wu, Weiping Wang 0005, Qixiang Ye |
Pattern Recognit. | 7 |
| 2022 | Dynamic Perception Framework for Fine-Grained RecognitionabstractFine-grained recognition poses the challenge of discriminating categories with only small subtle visual differences, which can be easily overwhelmed by diverse appearance within categories. Conventional approaches generally locate discriminative parts and then recognize the part-based features. However, we find that tuning the effective receptive field (ERF) of the network to the task plays the key role, which enables significant regions to contribute more to the output. Inspired by the receptive field stimulation mechanism of the visual cortex, we propose a Dynamic Perception framework as a solution. Our framework adapts the ERF by considering the image space and the kernel space simultaneously. In the image space, the Spatial Selective Sampling module is adopted to enlarge informative regions locally. In the kernel space, Spatial Selective Kernel convolution is introduced to adapt different kernel sizes for regions of interest and backgrounds by embedding spatial attention in the multi-path convolution. Extensive experiments on challenging benchmarks, including CUB-200-2011, FGVC-Aircraft, and Stanford Cars, demonstrate that our method yields a performance boost over the state-of-the-art methods. Yao Ding 0006, Zhenjun Han, Yanzhao Zhou, Yi Zhu 0004, Jie Chen 0001, Qixiang Ye, Jianbin Jiao |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2022 | Convex-Hull Feature Adaptation for Oriented and Densely Packed Object DetectionabstractDetecting oriented and densely packed objects is a challenging problem considering that the receptive field intersection between objects causes spatial feature aliasing. In this paper, we propose a convex-hull feature adaptation (CFA) approach, with the aim to configure convolutional features in accordance with irregular object layouts. CFA roots in the convex-hull feature representation, which defines a set of dynamically sampled feature points guided by the convex intersection over union (CIoU) to bound object extent. CFA pursues optimal feature assignment by constructing convex-hull sets and iteratively splitting positive or negative convex-hulls. By simultaneously considering overlapping convex-hulls and objects and penalizing convex-hulls shared by multiple objects, CFA defines a systematic way to adapt convolutional features on regular grids to objects of irregular shapes. Experiments on DOTA and SKU110K-R datasets show that CFA achieved new state-of-the-art performance for detecting oriented and densely packed objects. CFA also sets a solid baseline for convex polygon prediction on the MS COCO dataset defined for general object detection. Code is available athttps://github.com/SDL-GuoZonghao/BeyondBoundingBox. Zonghao Guo, Xiaosong Zhang 0004, Chang Liu 0047, Xiangyang Ji, Jianbin Jiao, Qixiang Ye |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2022 | Domain Contrast for Domain Adaptive Object DetectionabstractDespite of the substantial progress of visual object detection, models trained in one video domain often fail to generalize well to others due to the change of camera configurations, lighting conditions, and object person views. In this paper, we present Domain Contrast (DC), a simple yet effective approach inspired by contrastive learning for training domain adaptive detectors. DC is deduced from the error bound minimization perspective of a transferred model, and is implemented with cross-domain contrast loss which is plug-and-play. By minimizing cross-domain contrast loss, DC transfers detectors across domains while naturally alleviating the class imbalance issue in the target domain. DC can be applied at either image level or region level, consistently improving detectors’ discriminability while maintaining the transferability. Extensive experiments on commonly used benchmarks show that DC improves the baseline and state-of-the-art by significant margins, while demonstrating great potential for large domain divergence. Code is released athttps://github.com/PhoneSix/Domain-Contrast. Feng Liu 0050, Xiaosong Zhang 0004, Fang Wan 0001, Xiangyang Ji, Qixiang Ye |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | Disentangling Task-Oriented Representations for Unsupervised Domain AdaptationabstractUnsupervised domain adaptation (UDA) aims to address the domain-shift problem between a labeled source domain and an unlabeled target domain. Many efforts have been made to eliminate the mismatch between the distributions of training and testing data by learning domain-invariant representations. However, the learned representations are usually not task-oriented, i.e., being class-discriminative and domain-transferable simultaneously. This drawback limits the flexibility of UDA in complicated open-set tasks where no labels are shared between domains. In this paper, we break the concept of task-orientation into task-relevance and task-irrelevance, and propose a dynamic task-oriented disentangling network (DTDN) to learn disentangled representations in an end-to-end fashion for UDA. The dynamic disentangling network effectively disentangles data representations into two components: the task-relevant ones embedding critical information associated with the task across domains, and the task-irrelevant ones with the remaining non-transferable or disturbing information. These two components are regularized by a group of task-specific objective functions across domains. Such regularization explicitly encourages disentangling and avoids the use of generative models or decoders. Experiments in complicated, open-set scenarios (retrieval tasks) and empirical benchmarks (classification tasks) demonstrate that the proposed method captures rich disentangled information and achieves superior performance. Pingyang Dai, Peixian Chen, Qiong Wu 0012, Xiaopeng Hong, Qixiang Ye, Qi Tian 0001, Chia-Wen Lin, Rongrong Ji |
IEEE Trans. Image Process. | 5 |
| 2022 | Feature Calibration Network for Occluded Pedestrian DetectionabstractPedestrian detection in the wild remains a challenging problem especially for scenes containing serious occlusion. In this paper, we propose a novel feature learning method in the deep learning framework, referred to as Feature Calibration Network (FC-Net), to adaptively detect pedestrians under various occlusions. FC-Net is based on the observation that the visible parts of pedestrians are selective and decisive for detection, and is implemented as a self-paced feature learning framework with a self-activation (SA) module and a feature calibration (FC) module. In a new self-activated manner, FC-Net learns features which highlight the visible parts and suppress the occluded parts of pedestrians. The SA module estimates pedestrian activation maps by reusing classifier weights, without any additional parameter involved, therefore resulting in an extremely parsimony model to reinforce the semantics of features, while the FC module calibrates the convolutional features for adaptive pedestrian representation in both pixel-wise and region-based ways. Experiments on CityPersons and Caltech datasets demonstrate that FC-Net improves detection performance on occluded pedestrians up to 10% while maintaining excellent performance on non-occluded instances. Tianliang Zhang 0003, Qixiang Ye, Baochang Zhang 0001, Jianzhuang Liu, Xiaopeng Zhang 0008, Qi Tian 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2022 | Self-Guided Adaptation: Progressive Representation Alignment for Domain Adaptive Object DetectionabstractUnsupervised domain adaptation (UDA) has achieved unprecedented success in improving the cross-domain robustness of object detection models. However, existing UDA methods largely ignore the instantaneous data distribution and the sampling strategy during model learning, which could deteriorate the feature representation given large domain shift. In this work, we propose a Self-Guided Adaptation (SGA) model, targeting at aligning feature representation and transferring object detection models across domains while considering the instantaneous alignment difficulty. The core of SGA is to calculate “hardness” factors for sample pairs indicating domain distance in a kernel space. With the hardness factor, the proposed SGA adaptively indicates the importance of samples and assigns them different constrains. Indicated by these hardness factors, Self-Guided Progressive Sampling (SPS) is implemented in an “easy-to-hard” way during model adaptation. Using multi-stage convolutional features, SGA is further aggregated to fully align hierarchical representations of detection models. Extensive experiments on commonly-used benchmarks show that SGA improves the state-of-the-art methods with significant margins especially on large domain shift cases. Zongxian Li, Peixi Peng, Qixiang Ye, Shijian Lu, Tiejun Huang 0001, Yonghong Tian 0001 |
IEEE Trans. Multim. | 5 |
| 2022 | Filter Sketch for Network PruningabstractWe propose a novel network pruning approach by information preserving of pretrained network weights (filters). Network pruning with the information preserving is formulated as a matrix sketch problem, which is efficiently solved by the off-the-shelf frequent direction method. Our approach, referred to as FilterSketch, encodes the second-order information of pretrained weights, which enables the representation capacity of pruned networks to be recovered with a simple fine-tuning procedure. FilterSketch requires neither training from scratch nor data-driven iterative optimization, leading to a several-orders-of-magnitude reduction of time cost in the optimization of pruning. Experiments on CIFAR-10 show that FilterSketch reduces 63.3% of floating-point operations (FLOPs) and prunes 59.9% of network parameters with negligible accuracy cost for ResNet-110. On ILSVRC-2012, it reduces 45.5% of FLOPs and removes 43.0% of parameters with only 0.69% accuracy drop for ResNet-50. Our code and pruned models can be found at https://github.com/lmbxmu/FilterSketch. Mingbao Lin, Liujuan Cao, Qixiang Ye, Yonghong Tian 0001, Jianzhuang Liu, Qi Tian 0001, Rongrong Ji |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2022 | Network Pruning Using Adaptive Exemplar FiltersabstractPopular network pruning algorithms reduce redundant information by optimizing hand-crafted models, and may cause suboptimal performance and long time in selecting filters. We innovatively introduce adaptive exemplar filters to simplify the algorithm design, resulting in an automatic and efficient pruning approach called EPruner. Inspired by the face recognition community, we use a message-passing algorithm Affinity Propagation on the weight matrices to obtain an adaptive number of exemplars, which then act as the preserved filters. EPruner breaks the dependence on the training data in determining the "important" filters and allows the CPU implementation in seconds, an order of magnitude faster than GPU-based SOTAs. Moreover, we show that the weights of exemplars provide a better initialization for the fine-tuning. On VGGNet-16, EPruner achieves a 76.34%-FLOPs reduction by removing 88.80% parameters, with 0.06% accuracy improvement on CIFAR-10. In ResNet-152, EPruner achieves a 65.12%-FLOPs reduction by removing 64.18% parameters, with only 0.71% top-5 accuracy loss on ILSVRC-2012. Our code is available at https://github.com/lmbxmu/EPruner. Mingbao Lin, Rongrong Ji, Yan Wang 0059, Yongjian Wu 0001, Feiyue Huang, Qixiang Ye |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2022 | iffDetector: Inference-Aware Feature Filtering for Object DetectionabstractModern convolutional neural network (CNN)-based object detectors focus on feature configuration during training but often ignore feature optimization during inference. In this article, we propose a new feature optimization approach to enhance features and suppress background noise in both the training and inference stages. We introduce a generic inference-aware feature filtering (IFF) module that can be easily combined with existing detectors, resulting in our iffDetector. Unlike conventional open-loop feature calculation approaches without feedback, the proposed IFF module performs the closed-loop feature optimization by leveraging high-level semantics to enhance the convolutional features. By applying the Fourier transform to analyze our detector, we prove that the IFF module acts as a negative feedback that can theoretically guarantee the stability of the feature learning. IFF can be fused with CNN-based object detectors in a plug-and-play manner with little computational cost overhead. Experiments on the PASCAL VOC and MS COCO datasets demonstrate that our iffDetector consistently outperforms state-of-the-art methods with significant margins. Mingyuan Mao, Baochang Zhang 0001, Qixiang Ye, Wanquan Liu, David S. Doermann |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2022 | Part-Based Semantic Transform for Few-Shot Semantic SegmentationabstractFew-shot semantic segmentation remains an open problem for the lack of an effective method to handle the semantic misalignment between objects. In this article, we propose part-based semantic transform (PST) and target at aligning object semantics in support images with those in query images by semantic decomposition-and-match. The semantic decomposition process is implemented with prototype mixture models (PMMs), which use an expectation-maximization (EM) algorithm to decompose object semantics into multiple prototypes corresponding to object parts. The semantic match between prototypes is performed with a min-cost flow module, which encourages correct correspondence while depressing mismatches between object parts. With semantic decomposition-and-match, PST enforces the network's tolerance to objects' appearance and/or pose variation and facilities channelwise and spatial semantic activation of objects in query images. Extensive experiments on Pascal VOC and MS-COCO datasets show that PST significantly improves upon state-of-the-arts. In particular, on MS-COCO, it improves the performance of five-shot semantic segmentation by up to 7.79% with a moderate cost of inference speed and model size. Code for PST is released at https://github.com/Yang-Bob/PST. Boyu Yang 0002, Fang Wan 0001, Chang Liu 0047, Xiangyang Ji, Qixiang Ye |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2022 | Continuation Multiple Instance Learning for Weakly and Fully Supervised Object DetectionabstractWeakly supervised object detection (WSOD) is a challenging task that requires simultaneously learning object detectors and estimating object locations under the supervision of image category labels. Many WSOD methods that adopt multiple instance learning (MIL) have nonconvex objective functions and, therefore, are prone to get stuck in local minima (falsely localize object parts) while missing full object extent during training. In this article, we introduce classical continuation optimization into MIL, thereby creating continuation MIL (C-MIL) with the aim to alleviate the nonconvexity problem in a systematic way. To fulfill this purpose, we partition instances into class-related and spatially related subsets and approximate MIL's objective function with a series of smoothed objective functions defined within the subsets. We further propose a parametric strategy to implement continuation smooth functions, which enables C-MIL to be applied to instance selection tasks in a uniform manner. Optimizing smoothed loss functions prevents the training procedure from falling prematurely into local minima and facilities learning full object extent. Extensive experiments demonstrate the superiority of CMIL over conventional MIL methods. As a general instance selection method, C-MIL is also applied to supervised object detection to optimize anchors/features, improving the detection performance with a significant margin. Qixiang Ye, Fang Wan 0001, Chang Liu 0047, Qingming Huang, Xiangyang Ji |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2022 | Configurable Graph Reasoning for Visual Relationship DetectionabstractVisual commonsense knowledge has received growing attention in the reasoning of long-tailed visual relationships biased in terms of object and relation labels. Most current methods typically collect and utilize external knowledge for visual relationships by following the fixed reasoning path of {subject, object → predicate} to facilitate the recognition of infrequent relationships. However, the knowledge incorporation for such fixed multidependent path suffers from the data set biased and exponentially grown combinations of object and relation labels and ignores the semantic gap between commonsense knowledge and real scenes. To alleviate this, we propose configurable graph reasoning (CGR) to decompose the reasoning path of visual relationships and the incorporation of external knowledge, achieving configurable knowledge selection and personalized graph reasoning for each relation type in each image. Given a commonsense knowledge graph, CGR learns to match and retrieve knowledge for different subpaths and selectively compose the knowledge routed path. CGR adaptively configures the reasoning path based on the knowledge graph, bridges the semantic gap between the commonsense knowledge, and the real-world scenes and achieves better knowledge generalization. Extensive experiments show that CGR consistently outperforms previous state-of-the-art methods on several popular benchmarks and works well with different knowledge graphs. Detailed analyses demonstrated that CGR learned explainable and compelling configurations of reasoning paths. Yi Zhu 0004, Xiwen Liang, Bingqian Lin, Qixiang Ye, Jianbin Jiao, Liang Lin 0004, Xiaodan Liang |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2021 | Agreement-Discrepancy-Selection: Active Learning with Progressive Distribution AlignmentabstractIn active learning, the ignorance of aligning unlabeled samples' distribution with that of labeled samples hinders the model trained upon labeled samples from selecting informative unlabeled samples. In this paper, we propose an agreement-discrepancy-selection (ADS) approach, and target at unifying distribution alignment with sample selection by introducing adversarial classifiers to the convolutional neural network (CNN). Minimizing classifiers' prediction discrepancy (maximizing prediction agreement) drives learning CNN features to reduce the distribution bias of labeled and unlabeled samples, while maximizing classifiers' discrepancy highlights informative samples. Iterative optimization of agreement and discrepancy loss calibrated with an entropy function drives aligning sample distributions in a progressive fashion for effective active learning. Experiments on image classification and object detection tasks demonstrate that ADS is task-agnostic, while significantly outperforms the previous methods when the labeled sets are small. Mengying Fu, Tianning Yuan, Fang Wan 0001, Songcen Xu, Qixiang Ye |
AAAI | 5 |
| 2021 | Domain General Face Forgery Detection by Learning to WeightabstractIn this paper, we propose a domain-general model, termed learning-to-weight (LTW), that guarantees face detection performance across multiple domains, particularly the target domains that are never seen before. However, various face forgery methods cause complex and biased data distributions, making it challenging to detect fake faces in unseen domains. We argue that different faces contribute differently to a detection model trained on multiple domains, making the model likely to fit domain-specific biases. As such, we propose the LTW approach based on the meta-weight learning algorithm, which configures different weights for face images from different domains. The LTW network can balance the model's generalizability across multiple domains. Then, the meta-optimization calibrates the source domain's gradient enabling more discriminative features to be learned. The detection ability of the network is further improved by introducing an intra-class compact loss. Extensive experiments on several commonly used deepfake datasets to demonstrate the effectiveness of our method in detecting synthetic faces. Code and supplemental material are available at https://github.com/skJack/LTW. Ke Sun 0016, Hong Liu 0009, Qixiang Ye, Yue Gao 0002, Jianzhuang Liu, Ling Shao 0001, Rongrong Ji |
AAAI | 3 |
| 2021 | Nearest Neighbor Classifier Embedded Network for Active LearningabstractDeep neural networks (DNNs) have been widely applied to active learning. Despite of its effectiveness, the generalization ability of the discriminative classifier (the softmax classifier) is questionable when there is a significant distribution bias between the labeled set and the unlabeled set. In this paper, we attempt to replace the softmax classifier in deep neural network with a nearest neighbor classifier, considering its progressive generalization ability within the unknown sub-space. Our proposed active learning approach, termed nearest Neighbor Classifier Embedded network (NCE-Net), targets at reducing the risk of over-estimating unlabeled samples while improving the opportunity to query informative samples. NCE-Net is conceptually simple but surprisingly powerful, as justified from the perspective of the subset information, which defines a metric to quantify model generalization ability in active learning. Experimental results show that, with simple selection based on rejection or confusion confidence, NCE-Net improves state-of-the-arts on image classification and object detection tasks with significant margins. Fang Wan 0001, Tianning Yuan, Mengying Fu, Xiangyang Ji, Qingming Huang, Qixiang Ye |
AAAI | 6 |
| 2021 | Beyond Bounding-Box: Convex-Hull Feature Adaptation for Oriented and Densely Packed Object DetectionabstractDetecting oriented and densely packed objects remains challenging for spatial feature aliasing caused by the intersection of reception fields between objects. In this paper, we propose a convex-hull feature adaptation (CFA) approach for configuring convolutional features in accordance with oriented and densely packed object layouts. CFA is rooted in convex-hull feature representation, which defines a set of dynamically predicted feature points guided by the convex intersection over union (CIoU) to bound the extent of objects. CFA pursues optimal feature assignment by constructing convex-hull sets and dynamically splitting positive or negative convex-hulls. By simultaneously considering overlapping convex-hulls and objects and penalizing convex-hulls shared by multiple objects, CFA alleviates spatial feature aliasing towards optimal feature adaptation. Experiments on DOTA and SKU110K-R datasets show that CFA significantly outperforms the baseline approach, achieving new state-of-the-art detection performance. Code is available at github.com/SDL-GuoZonghao/BeyondBoundingBox. Zonghao Guo, Chang Liu 0042, Xiaosong Zhang 0004, Jianbin Jiao, Xiangyang Ji, Qixiang Ye |
CVPR | 6 |
| 2021 | Towards Compact CNNs via Collaborative CompressionabstractChannel pruning and tensor decomposition have received extensive attention in convolutional neural network compression. However, these two techniques are traditionally deployed in an isolated manner, leading to significant accuracy drop when pursuing high compression rates. In this paper, we propose a Collaborative Compression (CC) scheme, which joints channel pruning and tensor decomposition to compress CNN models by simultaneously learning the model sparsity and low-rankness. Specifically, we first investigate the compression sensitivity of each layer in the network, and then propose a Global Compression Rate Optimization that transforms the decision problem of compression rate into an optimization problem. After that, we propose multi-step heuristic compression to remove redundant compression units step-by-step, which fully considers the effect of the remaining compression space (i.e., unremoved compression units). Our method demonstrates superior performance gains over previous ones on various datasets and backbone architectures. For example, we achieve 52.9% FLOPs reduction by removing 48.4% parameters on ResNet-50 with only a Top-1 accuracy drop of 0.56% on ImageNet 2012. Shaohui Lin, Jianzhuang Liu, Qixiang Ye, Mengdi Wang 0001, Fei Chao 0001, Fan Yang 0016, Jincheng Ma, Qi Tian 0001, Rongrong Ji |
CVPR | 4 |
| 2021 | Beyond Max-Margin: Class Margin Equilibrium for Few-Shot Object DetectionabstractFew-shot object detection has made substantial progress by representing novel class objects using the feature representation learned upon a set of base class objects. However, an implicit contradiction between novel class classification and representation is unfortunately ignored. On the one hand, to achieve accurate novel class classification, the distributions of either two base classes must be far away from each other (max-margin). On the other hand, to precisely represent novel classes, the distributions of base classes should be close to each other to reduce the intra-class distance of novel classes (min-margin). In this paper, we propose a class margin equilibrium (CME) approach, with the aim to optimize both feature space partition and novel class reconstruction in a systematic way. CME first converts the few-shot detection problem to the few-shot classification problem by using a fully connected layer to decouple localization features. CME then reserves adequate margin space for novel classes by introducing simple-yet-effective class margin loss during feature learning. Finally, CME pursues margin equilibrium by disturbing the features of novel class instances in an adversarial min-max fashion. Experiments on Pascal VOC and MS-COCO datasets show that CME significantly improves upon two baseline detectors (up to 3 ~ 5% in average), achieving state-of-the-art performance. Code is available at https://github.com/BohaoLee/CME. Boyu Yang 0002, Chang Liu 0042, Feng Liu 0050, Rongrong Ji, Qixiang Ye |
CVPR | 6 |
| 2021 | Anti-Aliasing Semantic Reconstruction for Few-Shot Semantic SegmentationabstractEncouraging progress in few-shot semantic segmentation has been made by leveraging features learned upon base classes with sufficient training data to represent novel classes with few-shot examples. However, this feature sharing mechanism inevitably causes semantic aliasing between novel classes when they have similar compositions of semantic concepts. In this paper, we reformulate few-shot segmentation as a semantic reconstruction problem, and convert base class features into a series of basis vectors which span a class-level semantic space for novel class reconstruction. By introducing contrastive loss, we maximize the orthogonality of basis vectors while minimizing semantic aliasing between classes. Within the reconstructed representation space, we further suppress interference from other classes by projecting query features to the support vector for precise semantic activation. Our proposed approach, referred to as anti-aliasing semantic reconstruction (ASR), provides a systematic yet interpretable solution for few-shot learning problems. Extensive experiments on PASCAL VOC and MS COCO datasets show that ASR achieves strong results compared with the prior works. Code will be released at github.com/Bibkiller/ASR. Binghao Liu, Yao Ding 0006, Jianbin Jiao, Xiangyang Ji, Qixiang Ye |
CVPR | 5 |
| 2021 | Multiple Instance Active Learning for Object DetectionabstractDespite the substantial progress of active learning for image recognition, there still lacks an instance-level active learning method specified for object detection. In this paper, we propose Multiple Instance Active Object Detection (MI-AOD), to select the most informative images for detector training by observing instance-level uncertainty. MI-AOD defines an instance uncertainty learning module, which leverages the discrepancy of two adversarial instance classifiers trained on the labeled set to predict instance uncertainty of the unlabeled set. MI-AOD treats unlabeled images as instance bags and feature anchors in images as instances, and estimates the image uncertainty by re-weighting instances in a multiple instance learning (MIL) fashion. Iterative instance uncertainty learning and re-weighting facilitate suppressing noisy instances, toward bridging the gap between instance uncertainty and image-level uncertainty. Experiments validate that MI-AOD sets a solid baseline for instance-level active learning. On commonly used object detection datasets, MI-AOD outperforms state-of-the-art methods with significant margins, particularly when the labeled sets are small. Code is available at https://github.com/yuantn/MI-AOD. Tianning Yuan, Fang Wan 0001, Mengying Fu, Jianzhuang Liu, Songcen Xu, Xiangyang Ji, Qixiang Ye |
CVPR | 7 |
| 2021 | Video Enhancement Network Based on Max-Pooling and Hierarchical Feature FusionabstractIn this paper, we propose an efficient convolution neural network to enhance the quality of video compressed by HEVC standard. The model is composed of a max-pooling module and a hierarchical feature fusion module. The max-pooling module extracts feature from different scales and enlarges the receptive field of the model without stacking too many convolution layers. And the hierarchical feature fusion module accurately aligns features from different scales and fuses them efficiently. Two modules are applied in the proposed network, our model reconstructs compressed video frames with higher visual quality. Besides, the model is constructed in the full convolution network, thus it can adapt to videos in variable resolutions. The experiment results show that the proposed model outperforms existing models in the terms of PSNR under the same dataset. Honggang Qi, Jinwen Zan, Qixiang Ye, Guoqin Cui |
DCC | 5 |
| 2021 | Self-Motivated Communication Agent for Real-World Vision-Dialog NavigationabstractVision-Dialog Navigation (VDN) requires an agent to ask questions and navigate following the human responses to find target objects. Conventional approaches are only allowed to ask questions at predefined locations, which are built upon expensive dialogue annotations, and inconvenience the real-word human-robot communication and cooperation. In this paper, we propose a Self-Motivated Communication Agent (SCoA) that learns whether and what to communicate with human adaptively to acquire instructive information for realizing dialogue annotation-free navigation and enhancing the transferability in real-world unseen environment. Specifically, we introduce a whether-to-ask (WeTA) policy, together with uncertainty of which action to choose, to indicate whether the agent should ask a question. Then, a what-to-ask (WaTA) policy is proposed, in which, along with the oracle’s answers, the agent learns to score question candidates so as to pick up the most informative one for navigation, and meanwhile mimic oracle’s answering. Thus, the agent can navigate in a self-Q&A manner even in real-world environment where the human assistance is often unavailable. Through joint optimization of communication and navigation in a unified imitation learning and reinforcement learning framework, SCoA asks a question if necessary and obtains a hint for guiding the agent to move towards the target with less communication cost. Experiments on seen and unseen environments demonstrate that SCoA shows not only superior performance over existing baselines without dialog annotations, but also competing results compared with rich dialog annotations based counterparts. Yi Zhu 0004, Yue Weng, Fengda Zhu, Xiaodan Liang, Qixiang Ye, Yutong Lu, Jianbin Jiao |
ICCV | 5 |
| 2021 | Architecture Disentanglement for Deep Neural NetworksabstractUnderstanding the inner workings of deep neural networks (DNNs) is essential to provide trustworthy artificial intelligence techniques for practical applications. Existing studies typically involve linking semantic concepts to units or layers of DNNs, but fail to explain the inference process. In this paper, we introduce neural architecture disentanglement (NAD) to fill the gap. Specifically, NAD learns to disentangle a pre-trained DNN into sub-architectures according to independent tasks, forming information flows that describe the inference processes. We investigate whether, where, and how the disentanglement occurs through experiments conducted with handcrafted and automatically-searched network architectures, on both object-based and scene-based datasets. Based on the experimental results, we present three new findings that provide fresh insights into the inner logic of DNNs. First, DNNs can be divided into sub-architectures for independent tasks. Second, deeper layers do not always correspond to higher semantics. Third, the connection type in a DNN affects how the information flows across layers, leading to different disentanglement behaviors. With NAD, we further explain why DNNs sometimes give wrong predictions. Experimental results show that misclassified images have a high probability of being assigned to task sub-architectures similar to the correct ones. Our code is available at https://github.com/hujiecpp/NAD. Jie Hu 0018, Liujuan Cao, Qixiang Ye, Shengchuan Zhang, Ke Li 0015, Feiyue Huang, Ling Shao 0001, Rongrong Ji |
ICCV | 4 |
| 2021 | Occlude Them All: Occlusion-Aware Attention Network for Occluded Person Re-IDabstractPerson Re-Identification (ReID) has achieved remarkable performance along with the deep learning era. However, most approaches carry out ReID only based upon holistic pedestrian regions. In contrast, real-world scenarios involve occluded pedestrians, which provide partial visual appearances and destroy the ReID accuracy. A common strategy is to locate visible body parts by auxiliary model, which however suffers from significant domain gaps and data bias issues. To avoid such problematic models in occluded person ReID, we propose the Occlusion-Aware Mask Network (OAMN). In particular, we incorporate an attention-guided mask module, which requires guidance from labeled occlusion data. To this end, we propose a novel occlusion augmentation scheme that produces diverse and precisely labeled occlusion for any holistic dataset. The proposed scheme suits real-world scenarios better than existing schemes, which consider only limited types of occlusions. We also offer a novel occlusion unification scheme to tackle ambiguity information at the test phase. The above three components enable existing attention mechanisms to precisely capture body parts regardless of the occlusion. Comprehensive experiments on a variety of person ReID benchmarks demonstrate the superiority of OAMN over state-of-the-arts. Peixian Chen, Pingyang Dai, Jianzhuang Liu, Qixiang Ye, Mingliang Xu 0001, Qi'an Chen, Rongrong Ji |
ICCV | 5 |
| 2021 | TS-CAM: Token Semantic Coupled Attention Map for Weakly Supervised Object LocalizationabstractWeakly supervised object localization (WSOL) is a challenging problem when given image category labels but requires to learn object localization models. Optimizing a convolutional neural network (CNN) for classification tends to activate local discriminative regions while ignoring complete object extent, causing the partial activation issue. In this paper, we argue that partial activation is caused by the intrinsic characteristics of CNN, where the convolution operations produce local receptive fields and experience difficulty to capture long-range feature dependency among pixels. We introduce the token semantic coupled attention map (TS-CAM) to take full advantage of the self-attention mechanism in visual transformer for long-range dependency extraction. TS-CAM first splits an image into a sequence of patch tokens for spatial embedding, which produce attention maps of long-range visual dependency to avoid partial activation. TS-CAM then re-allocates category-related semantics for patch tokens, enabling each of them to be aware of object categories. TS-CAM finally couples the patch tokens with the semantic-agnostic attention map to achieve semantic-aware localization. Experiments on the ILSVRC/CUB-200-2011 datasets show that TS-CAM outperforms its CNN-CAM counterparts by 7.1%/27.1% for WSOL, achieving state-of-the-art performance. Code is available at https://github.com/vasgaowei/TS-CAM Wei Gao 0050, Fang Wan 0001, Xingjia Pan, Zhiliang Peng, Qi Tian 0001, Zhenjun Han, Bolei Zhou, Qixiang Ye |
ICCV | 8 |
| 2021 | Conformer: Local Features Coupling Global Representations for Visual RecognitionabstractWithin Convolutional Neural Network (CNN), the convolution operations are good at extracting local features but experience difficulty to capture global representations. Within visual transformer, the cascaded self-attention modules can capture long-distance feature dependencies but unfortunately deteriorate local feature details. In this paper, we propose a hybrid network structure, termed Conformer, to take advantage of convolutional operations and self-attention mechanisms for enhanced representation learning. Conformer roots in the Feature Coupling Unit (FCU), which fuses local features and global representations under different resolutions in an interactive fashion. Conformer adopts a concurrent structure so that local features and global representations are retained to the maximum extent. Experiments show that Conformer, under the comparable parameter complexity, outperforms the visual transformer (DeiT-B) by 2.3% on ImageNet. On MSCOCO, it outperforms ResNet-101 by 3.7% and 3.6% mAPs for object detection and instance segmentation, respectively, demonstrating the great potential to be a general backbone network. Code is available at github.com/pengzhiliang/Conformer. Zhiliang Peng, Shanzhi Gu, Lingxi Xie, Yaowei Wang 0001, Jianbin Jiao, Qixiang Ye |
ICCV | 7 |
| 2021 | Few-shot Learning for Multi-Modality TasksabstractRecent deep learning methods rely on a large amount of labeled data to achieve high performance. These methods may be impractical in some scenarios, where manual data annotation is costly or the samples of certain categories are scarce (e.g., tumor lesions, endangered animals and rare individual activities). When only limited annotated samples are available, these methods usually suffer from the overfitting problem severely, which degrades the performance significantly. In contrast, humans can recognize the objects in the images rapidly and correctly with their prior knowledge after exposed to only a few annotated samples. To simulate the learning schema of humans and relieve the reliance on the large-scale annotation benchmarks, researchers start shifting towards the few-shot learning problem: they try to learn a model to correctly recognize novel categories with only a few annotated samples. Jie Chen 0001, Qixiang Ye, Xiaoshan Yang, Shaohua Kevin Zhou, Xiaopeng Hong, Li Zhang 0040 |
ACM Multimedia | 2 |
| 2021 | Long-tailed Distribution AdaptationabstractRecognizing images with long-tailed distributions remains a challenging problem while there lacks an interpretable mechanism to solve this problem. In this study, we formulate Long-tailed recognition as Domain Adaption (LDA), by modeling the long-tailed distribution as an unbalanced domain and the general distribution as a balanced domain. Within the balanced domain, we propose to slack the generalization error bound, which is defined upon the empirical risks of unbalanced and balanced domains and the divergence between them. We propose to jointly optimize empirical risks of the unbalanced and balanced domains and approximate their domain divergence by intra-class and inter-class distances, with the aim to adapt models trained on the long-tailed distribution to general distributions in an interpretable way. Experiments on benchmark datasets for image recognition, object detection, and instance segmentation validate that our LDA approach, beyond its interpretability, achieves state-of-the-art performance. Zhiliang Peng, Zonghao Guo, Xiaosong Zhang 0004, Jianbin Jiao, Qixiang Ye |
ACM Multimedia | 6 |
| 2021 | MIGO-NAS: Towards Fast and Generalizable Neural Architecture SearchabstractNeural architecture search (NAS) has achieved unprecedented performance in various computer vision tasks. However, most existing NAS methods are defected in search efficiency and model generalizability. In this paper, we propose a novel NAS framework, termed MIGO-NAS, with the aim to guarantee the efficiency and generalizability in arbitrary search spaces. On the one hand, we formulate the search space as a multivariate probabilistic distribution, which is then optimized by a novel multivariate information-geometric optimization (MIGO). By approximating the distribution with a sampling, training, and testing pipeline, MIGO guarantees the memory efficiency, training efficiency, and search flexibility. Besides, MIGO is the first time to decrease the estimation error of natural gradient in multivariate distribution. On the other hand, for a set of specific constraints, the neural architectures are generated by a novel dynamic programming network generation (DPNG), which significantly reduces the training cost under various hardware environments. Experiments validate the advantages of our approach over existing methods by establishing a superior accuracy and efficiency i.e., 2.39 test error on CIFAR-10 benchmark and 21.7 on ImageNet benchmark, with only 1.5 GPU hours and 96 GPU hours for searching, respectively. Besides, the searched architectures can be well generalize to computer vision tasks including object detection and semantic segmentation, i.e., 25× FLOPs compression, with 6.4 mAP gain over Pascal VOC dataset, and 29.9× FLOPs compression, with only 1.41 percent performance drop over Cityscapes dataset. The code is publicly available. Xiawu Zheng, Rongrong Ji, Qiang Wang 0060, Baochang Zhang 0001, Jie Chen 0001, Qixiang Ye, Feiyue Huang, Yonghong Tian 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2021 | Context-aware network for RGB-D salient object detection
Fangfang Liang, Lijuan Duan, Wei Ma 0008, Yuanhua Qiao, Qixiang Ye |
Pattern Recognit. | 6 |
| 2021 | Discretization-aware architecture search
Yunjie Tian, Chang Liu 0042, Lingxi Xie, Jianbin Jiao, Qixiang Ye |
Pattern Recognit. | 5 |
| 2021 | Scale-Residual Learning Network for Scene Text DetectionabstractDetecting incidentally captured text in the wild remains an open problem due to challenging factors including unconstrained scenarios and large scale variation. In this paper, we establish a large-scale scene text detection dataset (LS-Text), containing 36, 000 images and 270, 783 text instances with various scales and complex scenarios, to promote the research of text detection. We propose a Scale-residual Learning Network (SLN) to deal with the scale variation problem in a progressive optimization manner. Specifically, we integrate both learnable feature concatenation and feature up-sampling operator. It can effectively eliminate the residuals between the outputs of SLN and ground-truth text instances by processing both the Feature Fusion Residuals (FFR) and the Scale Transformation Residuals (STR), simultaneously. By stacking multi-scale feature maps in a deep-to-shallow manner, SLN continuously optimizes feature representation by accumulating strong semantic information and rich texture details in a scale-residual learning way. Extensive experimental results on five challenging datasets demonstrate the state-of-the-art performance of the proposed SLN model, and the challenging aspects related to real-world scenarios of the proposed LS-Text dataset. Both the source code of SLN and the LS-Text dataset are available athttps://github.com/SLN-Text-Detection. Yuanqiang Cai, Chang Liu 0047, Peirui Cheng, Dawei Du, Libo Zhang 0001, Weiqiang Wang 0001, Qixiang Ye |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2021 | A Graphical Social Topology Model for RGB-D Multi-Person TrackingabstractTracking multiple persons is a challenging task especially when persons move in groups and occlude one another. Existing research have investigated the problems of group division and segmentation; however, lacking overall person-group topology modeling limits the ability to handle complex person and group dynamics. We propose a Graphical Social Topology (GST) model in the RGB-D data domain, and estimate object group dynamics by jointly modeling the group structure and states of persons using RGB-D topological representation. With our topology representation, moving persons are not only assigned to groups, but also dynamically connected with each other, which enables in-group individuals to be correctively associated and the cohesion of each group to be precisely modeled. Using the learned typical topology pattern and group online update modules, we infer the birth/death and merging/splitting of dynamic groups. With the GST model, the proposed multi-person tracker can naturally facilitate the occlusion problem by treating the occluded object and other in-group members as a whole, while leveraging overall state transition. Experiments on different RGB-D and RGB datasets confirm that the proposed multi-person tracker improves the state-of-the-arts. Shan Gao 0003, Qixiang Ye, Li Liu 0002, Arjan Kuijper, Xiangyang Ji |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | Harmonic Feature Activation for Few-Shot Semantic SegmentationabstractFew-shot semantic segmentation remains an open problem because limited support (training) images are insufficient to represent the diverse semantics within target categories. Conventional methods typically model a target category solely using information from the support image(s), resulting in incomplete semantic activation. In this paper, we propose a novel few-shot segmentation approach, termed harmonic feature activation (HFA), with the aim to implement dense support-to-query semantic transform by incorporating the features of both query and support images. HFA is formulated as a bilinear model, which takes charge of the pixel-wise dense correlation (bilinear feature activation) between query and support images in a systematic way. HFA incorporates a low-rank decomposition procedure, which speeds up bilinear feature activation with negligible performance cost. In addition, a semantic diffusion procedure is fused with HFA, which further improves the global harmony and local consistency of the feature activation. Extensive experiments on commonly used datasets (PASCAL VOC and MS COCO) show that HFA improves the state-of-the-arts with significant margins. Code is available at https://github.com/Bibikiller/HFA. Binghao Liu, Jianbin Jiao, Qixiang Ye |
IEEE Trans. Image Process. | 3 |
| 2021 | Adaptive Linear Span Network for Object Skeleton DetectionabstractConventional networks for object skeleton detection are usually hand-crafted. Despite the effectiveness, hand-crafted network architectures lack the theoretical basis and require intensive prior knowledge to implement representation complementarity for objects/parts in different granularity. In this paper, we propose an adaptive linear span network (AdaLSN), driven by neural architecture search (NAS), to automatically configure and integrate scale-aware features for object skeleton detection. AdaLSN is formulated with the theory of linear span, which provides one of the earliest explanations for multi-scale deep feature fusion. AdaLSN is materialized by defining a mixed unit-pyramid search space, which goes beyond many existing search spaces using unit-level or pyramid-level features. Within the mixed space, we apply genetic architecture search to jointly optimize unit-level operations and pyramid-level connections for adaptive feature space expansion. AdaLSN substantiates its versatility by achieving significantly higher accuracy and latency trade-off compared with the state-of-the-arts. It also demonstrates general applicability to image-to-mask tasks such as edge detection and road extraction. Code is available at https://github.com/sunsmarterjie/SDL-Skeletongithub.com/sunsmarterjie/SDL-Skeleton. Chang Liu 0047, Yunjie Tian, Zhiwen Chen 0002, Jianbin Jiao, Qixiang Ye |
IEEE Trans. Image Process. | 5 |
| 2021 | SRN: Side-Output Residual Network for Object Reflection Symmetry Detection and BeyondabstractThis article establishes a baseline for object reflection symmetry detection in natural images by releasing a new benchmark named Sym-PASCAL and proposing an end-to-end deep learning approach for reflection symmetry. Sym-PASCAL spans challenges of multiobjects, object diversity, part invisibility, and clustered backgrounds, which is far beyond those in existing data sets. The end-to-end deep learning approach, referred to as a side-output residual network (SRN), leverages the output residual units (RUs) to fit the errors between the symmetry ground truth and the side outputs of multiple stages of a trunk network. By cascading RUs from deep to shallow, SRN exploits the "flow" of errors along multiple stages to effectively matching object symmetry at different scales and suppress the clustered backgrounds. SRN is interpreted as a boosting-like algorithm, which assembles features using RUs during network forward and backward propagations. SRN is further upgraded to a multitask SRN (MT-SRN) for joint symmetry and edge detection, demonstrating its generality to image-to-mask learning tasks. Experimental results verify that the Sym-PASCAL benchmark is challenging related to real-world images, SRN achieves state-of-the-art performance, and MT-SRN has the capability to simultaneously predict edge and symmetry mask without loss of performance. Wei Ke 0003, Jie Chen 0001, Jianbin Jiao, Guoying Zhao 0001, Qixiang Ye |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2020 | SPSTracker: Sub-Peak Suppression of Response Map for Robust Object TrackingabstractModern visual trackers usually construct online learning models under the assumption that the feature response has a Gaussian distribution with target-centered peak response. Nevertheless, such an assumption is implausible when there is progressive interference from other targets and/or background noise, which produce sub-peaks on the tracking response map and cause model drift. In this paper, we propose a rectified online learning approach for sub-peak response suppression and peak response enforcement and target at handling progressive interference in a systematic way. Our approach, referred to as SPSTracker, applies simple-yet-efficient Peak Response Pooling (PRP) to aggregate and align discriminative features, as well as leveraging a Boundary Response Truncation (BRT) to reduce the variance of feature response. By fusing with multi-scale features, SPSTracker aggregates the response distribution of multiple sub-peaks to a single maximum peak, which enforces the discriminative capability of features for robust object tracking. Experiments on the OTB, NFS and VOT2018 benchmarks demonstrate that SPSTrack outperforms the state-of-the-art real-time trackers with significant margins1 Qintao Hu, Yao Mao, Jianlin Zhang 0001, Qixiang Ye |
AAAI | 6 |
| 2020 | Video Cloze Procedure for Self-Supervised Spatio-Temporal LearningabstractWe propose a novel self-supervised method, referred to as Video Cloze Procedure (VCP), to learn rich spatial-temporal representations. VCP first generates “blanks” by withholding video clips and then creates “options” by applying spatio-temporal operations on the withheld clips. Finally, it fills the blanks with “options” and learns representations by predicting the categories of operations applied on the clips. VCP can act as either a proxy task or a target task in self-supervised learning. As a proxy task, it converts rich self-supervised representations into video clip operations (options), which enhances the flexibility and reduces the complexity of representation learning. As a target task, it can assess learned representation models in a uniform and interpretable manner. With VCP, we train spatial-temporal representation models (3D-CNNs) and apply such models on action recognition and video retrieval tasks. Experiments on commonly used benchmarks show that the trained models outperform the state-of-the-art self-supervised models with significant margins. Dezhao Luo, Chang Liu 0042, Yu Zhou 0015, Dongbao Yang, Can Ma, Qixiang Ye, Weiping Wang 0005 |
AAAI | 6 |
| 2020 | Multiple Anchor Learning for Visual Object DetectionabstractClassification and localization are two pillars of visual object detectors. However, in CNN-based detectors, these two modules are usually optimized under a fixed set of candidate (or anchor) bounding boxes. This configuration significantly limits the possibility to jointly optimize classification and localization. In this paper, we propose a Multiple Instance Learning (MIL) approach that selects anchors and jointly optimizes the two modules of a CNN-based object detector. Our approach, referred to as Multiple Anchor Learning (MAL), constructs anchor bags and selects the most representative anchors from each bag. Such an iterative selection process is potentially NP-hard to optimize. To address this issue, we solve MAL by repetitively depressing the confidence of selected anchors by perturbing their corresponding features. In an adversarial selection-depression manner, MAL not only pursues optimal solutions but also fully leverages multiple anchors/features to learn a detection model. Experiments show that MAL improves the baseline RetinaNet with significant margins on the commonly used MS-COCO object detection benchmark and achieves new state-of-the-art detection performance compared with recent methods. Wei Ke 0003, Tianliang Zhang 0003, Zeyi Huang, Qixiang Ye, Jianzhuang Liu |
CVPR | 4 |
| 2020 | Video Playback Rate Perception for Self-Supervised Spatio-Temporal Representation LearningabstractIn self-supervised spatio-temporal representation learning, the temporal resolution and long-short term characteristics are not yet fully explored, which limits representation capabilities of learned models. In this paper, we propose a novel self-supervised method, referred to as video Playback Rate Perception (PRP), to learn spatio-temporal representation in a simple-yet-effective way. PRP roots in a dilated sampling strategy, which produces self-supervision signals about video playback rates for representation model learning. PRP is implemented with a feature encoder, a classification module, and a reconstructing decoder, to achieve spatio-temporal semantic retention in a collaborative discrimination-generation manner. The discriminative perception model follows a feature encoder to prefer perceiving low temporal resolution and long-term representation by classifying fast-forward rates. The generative perception model acts as a feature decoder to focus on comprehending high temporal resolution and short-term representation by introducing a motion-attention mechanism. PRP is applied on typical video target tasks including action recognition and video retrieval. Experiments show that PRP outperforms state-of-the-art self-supervised models with significant margins. Code is available at github.com/yuanyao366/PRP. Chang Liu 0042, Dezhao Luo, Yu Zhou 0015, Qixiang Ye |
CVPR | 5 |
| 2020 | AD-Cluster: Augmented Discriminative Clustering for Domain Adaptive Person Re-IdentificationabstractDomain adaptive person re-identification (re-ID) is a challenging task, especially when person identities in target domains are unknown. Existing methods attempt to address this challenge by transferring image styles or aligning feature distributions across domains, whereas the rich unlabeled samples in target domains are not sufficiently exploited. This paper presents a novel augmented discriminative clustering (AD-Cluster) technique that estimates and augments person clusters in target domains and enforces the discrimination ability of re-ID models with the augmented clusters. AD-Cluster is trained by iterative density-based clustering, adaptive sample augmentation, and discriminative feature learning. It learns an image generator and a feature encoder which aim to maximize the intra-cluster diversity in the sample space and minimize the intra-cluster distance in the feature space in an adversarial min-max manner. Finally, AD-Cluster increases the diversity of sample clusters and improves the discrimination capability of re-ID models greatly. Extensive experiments over Market-1501 and DukeMTMC-reID show that AD-Cluster outperforms the state-of-the-art with large margins. Yunpeng Zhai, Shijian Lu, Qixiang Ye, Xuebo Shan, Jie Chen 0001, Rongrong Ji, Yonghong Tian 0001 |
CVPR | 3 |
| 2020 | Rethinking Performance Estimation in Neural Architecture SearchabstractNeural architecture search (NAS) remains a challenging problem, which is attributed to the indispensable and time-consuming component of performance estimation (PE). In this paper, we provide a novel yet systematic rethinking of PE in a resource constrained regime, termed budgeted PE (BPE), which precisely and effectively estimates the performance of an architecture sampled from an architecture space. Since searching an optimal BPE is extremely time-consuming as it requires to train a large number of networks for evaluation, we propose a Minimum Importance Pruning (MIP) approach. Given a dataset and a BPE search space, MIP estimates the importance of hyper-parameters using random forest and subsequently prunes the minimum one from the next iteration. In this way, MIP effectively prunes less important hyper-parameters to allocate more computational resource on more important ones, thus achieving an effective exploration. By combining BPE with various search algorithms including reinforcement learning, evolution algorithm, random search, and differentiable architecture search, we achieve 1, 000× of NAS speed up with a negligible performance drop comparing to the SOTA. Xiawu Zheng, Rongrong Ji, Qiang Wang 0060, Qixiang Ye, Zhenguo Li, Yonghong Tian 0001, Qi Tian 0001 |
CVPR | 4 |
| 2020 | Cogradient Descent for Bilinear OptimizationabstractConventional learning methods simplify the bilinear model by regarding two intrinsically coupled factors independently, which degrades the optimization procedure. One reason lies in the insufficient training due to the asynchronous gradient descent, which results in vanishing gradients for the coupled variables. In this paper, we introduce a Cogradient Descent algorithm (CoGD) to address the bilinear problem, based on a theoretical framework to coordinate the gradient of hidden variables via a projection function. We solve one variable by considering its coupling relationship with the other, leading to a synchronous gradient descent to facilitate the optimization procedure. Our algorithm is applied to solve problems with one variable under the sparsity constraint, which is widely used in the learning paradigm. We validate our CoGD considering an extensive set of applications including image reconstruction, inpainting, and network pruning. Experiments show that it improves the state-of-the-art by a significant margin. Lian Zhuo, Baochang Zhang 0001, Linlin Yang 0001, Qixiang Ye, David S. Doermann, Rongrong Ji, Guodong Guo |
CVPR | 5 |
| 2020 | API-Net: Robust Generative Classifier via a Single Discriminator
Xinshuai Dong, Hong Liu 0009, Rongrong Ji, Liujuan Cao, Qixiang Ye, Jianzhuang Liu, Qi Tian 0001 |
ECCV (13) | 5 |
| 2020 | Component Divide-and-Conquer for Real-World Image Super-Resolution
Pengxu Wei, Ziwei Xie, Hannan Lu, Zongyuan Zhan, Qixiang Ye, Wangmeng Zuo, Liang Lin 0004 |
ECCV (8) | 5 |
| 2020 | Prototype Mixture Models for Few-Shot Semantic Segmentation
Boyu Yang 0002, Chang Liu 0042, Jianbin Jiao, Qixiang Ye |
ECCV (8) | 5 |
| 2020 | Multiple Expert Brainstorming for Domain Adaptive Person Re-Identification
Yunpeng Zhai, Qixiang Ye, Shijian Lu, Mengxi Jia, Rongrong Ji, Yonghong Tian 0001 |
ECCV (7) | 2 |
| 2020 | Progressive Cluster Purification for Unsupervised Feature LearningabstractIn unsupervised feature learning, sample specificity based methods ignore the inter-class information, which deteriorates the discriminative capability of representation models. Clustering based methods are error-prone to explore the complete class boundary information due to the inevitable class inconsistent samples in each cluster. In this work, we propose a novel clustering based method, which, by iteratively excluding class inconsistent samples during progressive cluster formation, alleviates the impact of noise samples in a simple-yet-effective manner. Our approach, referred to as Progressive Cluster Purification (PCP), implements progressive clustering by gradually reducing the number of clusters during training, while the sizes of clusters continuously expand consistently with the growth of model representation capability. With a well-designed cluster purification mechanism, it further purifies clusters by filtering noise samples which facilitate the subsequent feature learning by utilizing the refined clusters as pseudo-labels. Experiments on commonly used benchmarks demonstrate that the proposed PCP improves baseline method with significant margins. Our code will be available at https://github.com/zhangyifei0115/PCP. Yifei Zhang 0005, Chang Liu 0042, Yu Zhou 0015, Wei Wang 0315, Weiping Wang 0005, Qixiang Ye |
ICPR | 6 |
| 2020 | Scale Match for Tiny Person DetectionabstractVisual object detection has achieved unprecedented advance with the rise of deep convolutional neural networks. However, detecting tiny objects (for example tiny persons less than 20 pixels) in large-scale images remains not well investigated. The extremely small objects raise a grand challenge about feature representation while the massive and complex backgrounds aggregate the risk of false alarms. In this paper, we introduce a new benchmark, referred to as TinyPerson, opening up a promising direction for tiny object detection in a long distance and with massive backgrounds. We experimentally find that the scale mismatch between the dataset for network pre-training and the dataset for detector learning could deteriorate the feature representation and the detectors. Accordingly, we propose a simple yet effective Scale Match approach to align the object scales between the two datasets for favorable tiny-object representation. Experiments show the significant performance gain of our proposed approach over state-of-the-art detectors, and the challenging aspects of TinyPerson related to real-world scenarios. The TinyPerson benchmark and the code for our approach will be publicly available1. Xuehui Yu, Yuqi Gong, Qixiang Ye, Zhenjun Han |
WACV | 4 |
| 2020 | Image captioning via semantic element embedding
Xiaodan Zhang 0003, Shengfeng He, Xinhang Song, Rynson W. H. Lau, Jianbin Jiao, Qixiang Ye |
Neurocomputing | 6 |
| 2020 | IOS-Net: An inside-to-outside supervision network for scale robust text detection in the wild
Yuanqiang Cai, Weiqiang Wang 0001, Qixiang Ye |
Pattern Recognit. | 4 |
| 2020 | CoCNN: RGB-D deep fusion for stereoscopic salient object detection
Fangfang Liang, Lijuan Duan, Wei Ma 0008, Yuanhua Qiao, Zhi Cai, Qixiang Ye |
Pattern Recognit. | 7 |
| 2020 | Adaptive Discriminative Deep Correlation Filter for Visual Object TrackingabstractCorrelation filter trackers building on deep convolution neural networks (CNNs) contribute efficient visual object trackers but remain challenged with severe target appearance variations. The reason for this is that CNNs trained for image classification tasks are less discriminative to the dynamic variations of targets and backgrounds. In this paper, we propose an adaptive discriminative deep correlation filter (adaDDCF), which, by incorporating discriminative feature fine-tuning with adaptive appearance modeling, pursues stable object tracking in complex backgrounds. In adaDDCF, a convolutional Fisher discriminative analysis (FDA) layer is implemented for positive and negative instance mining and scene-specific feature learning. A correlation layer is then embedded to learn the correlation response of consecutive frames for target appearance modeling. With an online learning procedure using forward-backward propagation, the FDA layer and the correlation layer are effectively coupled, leading to effective and discriminative fine-tuning for the proposed tracker, which consequently alleviates the target drifting problem. Extensive experiments on the challenging benchmarks OTB2013, OTB2015, and OTB50 demonstrate that the proposed adaDDCF tracker outperforms many state-of-the-art trackers. Zhenjun Han, Qixiang Ye |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2020 | Progressive Latent Models for Self-Learning Scene-Specific Pedestrian DetectorsabstractThe performance of offline learned pedestrian detectors significantly drops when they are applied to video scenes of various camera views, occlusions, and background structures. Learning a detector for each video scene can avoid the performance drop but it requires repetitive human effort on data annotation. In this paper, a self-learning approach is proposed, toward specifying a pedestrian detector for each video scene without any human annotation involved. Object locations in video frames are treated as latent variables and a progressive latent model (PLM) is proposed to solve such latent variables. The PLM is deployed as components of object discovery, object enforcement, and label propagation, which are used to learn the object locations in a progressive manner. With the difference of convex (DC) objective functions, PLM is optimized by a concave-convex programming algorithm. With specified network branches and loss functions, PLM is integrated with deep feature learning and optimized in an end-to-end manner. From the perspectives of convex regularization and error rate estimation, detailed optimization analysis and learning stability analysis of the proposed PLM are provided. The extensive experiments demonstrate that even without annotation involved the proposed self-learning approach outperforms weakly supervised learning approaches, while achieving comparable performance with transfer learning approaches. Qixiang Ye, Tianliang Zhang 0003, Wei Ke 0003 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2020 | CircleNet: Reciprocating Feature Adaptation for Robust Pedestrian DetectionabstractPedestrian detection in the wild remains a challenging problem especially when the scene contains significant occlusion and/or low resolution of the pedestrians to be detected. Existing methods are unable to adapt to these difficult cases while maintaining acceptable performance. In this paper we propose a novel feature learning model, referred to as CircleNet, to achieve feature adaptation by mimicking the process humans looking at low resolution and occluded objects: focusing on it again, at a finer scale, if the object can not be identified clearly for the first time. CircleNet is implemented as a set of feature pyramids and uses weight sharing path augmentation for better feature fusion. It targets at reciprocating feature adaptation and iterative object detection using multiple top-down and bottom-up pathways. To take full advantage of the feature adaptation capability in CircleNet, we design an instance decomposition training strategy to focus on detecting pedestrian instances of various resolutions and different occlusion levels in each cycle. Specifically, CircleNet implements feature ensemble with the idea of hard negative boosting in an end-to-end manner. Experiments on two pedestrian detection datasets, Caltech and CityPersons, show that CircleNet improves the performance of occluded and low-resolution pedestrians with significant margins while maintaining good performance on normal instances. Tianliang Zhang 0003, Zhenjun Han, Huijuan Xu 0001, Baochang Zhang 0001, Qixiang Ye |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2019 | Calibrated Stochastic Gradient Descent for Convolutional Neural NetworksabstractIn stochastic gradient descent (SGD) and its variants, the optimized gradient estimators may be as expensive to compute as the true gradient in many scenarios. This paper introduces a calibrated stochastic gradient descent (CSGD) algorithm for deep neural network optimization. A theorem is developed to prove that an unbiased estimator for the network variables can be obtained in a probabilistic way based on the Lipschitz hypothesis. Our work is significantly distinct from existing gradient optimization methods, by providing a theoretical framework for unbiased variable estimation in the deep learning paradigm to optimize the model parameter calculation. In particular, we develop a generic gradient calibration layer which can be easily used to build convolutional neural networks (CNNs). Experimental results demonstrate that CNNs with our CSGD optimization scheme can improve the stateof-the-art performance for natural image classification, digit recognition, ImageNet object classification, and object detection tasks. This work opens new research directions for developing more efficient SGD updates and analyzing the backpropagation algorithm. Lian Zhuo, Baochang Zhang 0001, Chen Chen 0001, Qixiang Ye, Jianzhuang Liu, David S. Doermann |
AAAI | 4 |
| 2019 | Towards Optimal Structured CNN Pruning via Generative Adversarial LearningabstractStructured pruning of filters or neurons has received increased focus for compressing convolutional neural networks. Most existing methods rely on multi-stage optimizations in a layer-wise manner for iteratively pruning and retraining which may not be optimal and may be computation intensive. Besides, these methods are designed for pruning a specific structure, such as filter or block structures without jointly pruning heterogeneous structures. In this paper, we propose an effective structured pruning approach that jointly prunes filters as well as other structures in an end-to-end manner. To accomplish this, we first introduce a soft mask to scale the output of these structures by defining a new objective function with sparsity regularization to align the output of baseline and network with this mask. We then effectively solve the optimization problem by generative adversarial learning (GAL), which learns a sparse soft mask in a label-free and an end-to-end manner. By forcing more scale factors in the soft mask to zero, the fast iterative shrinkage-thresholding algorithm (FISTA) can be leveraged to fast and reliably remove the corresponding structures. Extensive experiments demonstrate the effectiveness of GAL on different datasets, including MNIST, CIFAR-10 and ImageNet ILSVRC 2012. For example, on ImageNet ILSVRC 2012, the pruned ResNet-50 achieves 10.88% Top-5 error and results in a factor of 3.7x speedup. This significantly outperforms state-of-the-art methods. Shaohui Lin, Rongrong Ji, Chenqian Yan, Baochang Zhang 0001, Liujuan Cao, Qixiang Ye, Feiyue Huang, David S. Doermann |
CVPR | 6 |
| 2019 | Orthogonal Decomposition Network for Pixel-Wise Binary ClassificationabstractThe weight sharing scheme and spatial pooling operations in Convolutional Neural Networks (CNNs) introduce semantic correlation to neighboring pixels on feature maps and therefore deteriorate their pixel-wise classification performance. In this paper, we implement an Orthogonal Decomposition Unit (ODU) that transforms a convolutional feature map into orthogonal bases targeting at de-correlating neighboring pixels on convolutional features. In theory, complete orthogonal decomposition produces orthogonal bases which can perfectly reconstruct any binary mask (ground-truth). In practice, we further design incomplete orthogonal decomposition focusing on de-correlating local patches which balances the reconstruction performance and computational cost. Fully Convolutional Networks (FCNs) implemented with ODUs, referred to as Orthogonal Decomposition Networks (ODNs), learn de-correlated and complementary convolutional features and fuse such features in a pixel-wise selective manner. Over pixel-wise binary classification tasks for two-dimensional image processing, specifically skeleton detection, edge detection, and saliency detection, and one-dimensional keypoint detection, specifically S-wave arrival time detection for earthquake localization, ODNs consistently improves the state-of-the-arts with significant margins. Chang Liu 0042, Fang Wan 0001, Wei Ke 0003, Zhuowei Xiao, Xiaosong Zhang 0004, Qixiang Ye |
CVPR | 7 |
| 2019 | SIXray: A Large-Scale Security Inspection X-Ray Benchmark for Prohibited Item Discovery in Overlapping ImagesabstractIn this paper, we present a large-scale dataset and establish a baseline for prohibited item discovery in Security Inspection X-ray images. Our dataset, named SIXray, consists of 1,059,231 X-ray images, in which 6 classes of 8,929 prohibited items are manually annotated. It raises a brand new challenge of overlapping image data, meanwhile shares the same properties with existing datasets, including complex yet meaningless contexts and class imbalance. We propose an approach named class-balanced hierarchical refinement (CHR) to deal with these difficulties. CHR assumes that each input image is sampled from a mixture distribution, and that deep networks require an iterative process to infer image contents accurately. To accelerate, we insert reversed connections to different network backbones, delivering high-level visual cues to assist mid-level features. In addition, a class-balanced loss function is designed to maximally alleviate the noise introduced by easy negative samples. We evaluate CHR on SIXray with different ratios of positive/negative samples. Compared to the baselines, CHR enjoys a better ability of discriminating objects especially using mid-level features, which offers the possibility of using a weakly-supervised approach towards accurate object localization. In particular, the advantage of CHR is more significant in the scenarios with fewer positive training samples, which demonstrates its potential application in real-world security inspection. Caijing Miao, Lingxi Xie, Fang Wan 0001, Chi Su, Hongye Liu, Jianbin Jiao, Qixiang Ye |
CVPR | 7 |
| 2019 | C-MIL: Continuation Multiple Instance Learning for Weakly Supervised Object DetectionabstractWeakly supervised object detection (WSOD) is a challenging task when provided with image category supervision but required to simultaneously learn object locations and object detectors. Many WSOD approaches adopt multiple instance learning (MIL) and have non-convex loss functions which are prone to get stuck into local minima (falsely localize object parts) while missing full object extent during training. In this paper, we introduce a continuation optimization method into MIL and thereby creating continuation multiple instance learning (C-MIL), with the intention of alleviating the non-convexity problem in a systematic way. We partition instances into spatially related and class related subsets, and approximate the original loss function with a series of smoothed loss functions defined within the subsets. Optimizing smoothed loss functions prevents the training procedure falling prematurely into local minima and facilitates the discovery of Stable Semantic Extremal Regions (SSERs) which indicate full object extent. On the PASCAL VOC 2007 and 2012 datasets, C-MIL improves the state-of-the-art of weakly supervised object detection and weakly supervised object localization with large margins. Fang Wan 0001, Chang Liu 0042, Wei Ke 0003, Xiangyang Ji, Jianbin Jiao, Qixiang Ye |
CVPR | 6 |
| 2019 | Learning Instance Activation Maps for Weakly Supervised Instance SegmentationabstractDiscriminative region responses residing inside an object instance can be extracted from networks trained with image-level label supervision. However, learning the full extent of pixel-level instance response in a weakly supervised manner remains unexplored. In this work, we tackle this challenging problem by using a novel instance extent filling approach. We first design a process to selectively collect pseudo supervision from noisy segment proposals obtained with previously published techniques. The pseudo supervision is used to learn a differentiable filling module that predicts a class-agnostic activation map for each instance given the image and an incomplete region response. We refer to the above maps as Instance Activation Maps (IAMs), which provide a fine-grained instance-level representation and allow instance masks to be extracted by lightweight CRF. Extensive experiments on the PASCAL VOC12 dataset show that our approach beats the state-of-the-art weakly supervised instance segmentation methods by a significant margin and increases the inference speed by an order of magnitude. Our method also generalizes well across domains and to unseen object categories. Without fine-tuning for the specific tasks, our model trained on VOC12 dataset (20 classes) obtains top performance for weakly supervised object localization on the CUB dataset (200 classes) and achieves competitive results on three widely used salient object detection benchmarks. Yi Zhu 0004, Yanzhao Zhou, Huijuan Xu 0001, Qixiang Ye, David S. Doermann, Jianbin Jiao |
CVPR | 4 |
| 2019 | Selective Sparse Sampling for Fine-Grained Image RecognitionabstractFine-grained recognition poses the unique challenge of capturing subtle inter-class differences under considerable intra-class variances (e.g., beaks for bird species). Conventional approaches crop local regions and learn detailed representation from those regions, but suffer from the fixed number of parts and missing of surrounding context. In this paper, we propose a simple yet effective framework, called Selective Sparse Sampling, to capture diverse and fine-grained details. The framework is implemented using Convolutional Neural Networks, referred to as Selective Sparse Sampling Networks (S3Ns). With image-level supervision, S3Ns collect peaks, i.e., local maximums, from class response maps to estimate informative, receptive fields and learn a set of sparse attention for capturing fine-detailed visual evidence as well as preserving context. The evidence is selectively sampled to extract discriminative and complementary features, which significantly enrich the learned representation and guide the network to discover more subtle cues. Extensive experiments and ablation studies show that the proposed method consistently outperforms the state-of-the-art methods on challenging benchmarks including CUB-200-2011, FGVC-Aircraft, and Stanford Cars. Yao Ding 0006, Yanzhao Zhou, Yi Zhu 0004, Qixiang Ye, Jianbin Jiao |
ICCV | 4 |
| 2019 | DANet: Divergent Activation for Weakly Supervised Object LocalizationabstractWeakly supervised object localization remains a challenge when learning object localization models from image category labels. Optimizing image classification tends to activate object parts and ignore the full object extent, while expanding object parts into full object extent could deteriorate the performance of image classification. In this paper, we propose a divergent activation (DA) approach, and target at learning complementary and discriminative visual patterns for image classification and weakly supervised object localization from the perspective of discrepancy. To this end, we design hierarchical divergent activation (HDA), which leverages the semantic discrepancy to spread feature activation, implicitly. We also propose discrepant divergent activation (DDA), which pursues object extent by learning mutually exclusive visual patterns, explicitly. Deep networks implemented with HDA and DDA, referred to as DANets, diverge and fuse discrepant yet discriminative features for image classification and object localization in an end-to-end manner. Experiments validate that DANets advance the performance of object localization while maintaining high performance of image classification on CUB-200 and ILSVRC datasets. Haolan Xue, Chang Liu 0042, Fang Wan 0001, Jianbin Jiao, Xiangyang Ji, Qixiang Ye |
ICCV | 6 |
| 2019 | Information Competing Process for Learning Diversified RepresentationsabstractLearning representations with diversified information remains as an open problem. Towards learning diversified representations, a new approach, termed Information Competing Process (ICP), is proposed in this paper. Aiming to enrich the information carried by feature representations, ICP separates a representation into two parts with different mutual information constraints. The separated parts are forced to accomplish the downstream task independently in a competitive environment which prevents the two parts from learning what each other learned for the downstream task. Such competing parts are then combined synergistically to complete the task. By fusing representation parts learned competitively under different conditions, ICP facilitates obtaining diversified representations which contain rich information. Experiments on image classification and image reconstruction tasks demonstrate the great potential of ICP to learn discriminative and disentangled representations in both supervised and self-supervised learning settings. Jie Hu 0018, Rongrong Ji, Shengchuan Zhang, Xiaoshuai Sun, Qixiang Ye, Chia-Wen Lin, Qi Tian 0001 |
NeurIPS | 5 |
| 2019 | FreeAnchor: Learning to Match Anchors for Visual Object DetectionabstractModern CNN-based object detectors assign anchors for ground-truth objects under the restriction of object-anchor Intersection-over-Unit (IoU). In this study, we propose a learning-to-match approach to break IoU restriction, allowing objects to match anchors in a flexible manner. Our approach, referred to as FreeAnchor, updates hand-crafted anchor assignment to "free" anchor matching by formulating detector training as a maximum likelihood estimation (MLE) procedure. FreeAnchor targets at learning features which best explain a class of objects in terms of both classification and localization. FreeAnchor is implemented by optimizing detection customized likelihood and can be fused with CNN-based detectors in a plug-and-play manner. Experiments on MS-COCO demonstrate that FreeAnchor consistently outperforms the counterparts with significant margins. Xiaosong Zhang 0004, Fang Wan 0001, Chang Liu 0042, Rongrong Ji, Qixiang Ye |
NeurIPS | 5 |
| 2019 | Starts Better and Ends Better: A Target Adaptive Image Signature TrackerabstractCorrelation filter (CF) trackers have achieved outstanding performance in visual object tracking tasks, in which the cosine mask plays an essential role in alleviating boundary effects caused by the circular assumption. However, the cosine mask imposes a larger weight on its center position, which greatly affects CF trackers, that is, their performance will drop significantly if a bad starting point happens to occur. To address the above issue, we propose a target adaptive image signature (TaiS) model to refine the starting point in each frame for CF trackers. Specifically, we incorporate the target priori into the image signature to build a target-specific saliency map, and iteratively refine the starting point with a closed-form solution during the tracking process. As a result, our TaiS is able to find a better starting point close to the center of targets; more importantly, it is independent of specific CF trackers and can efficiently improve their performance. Experiments on two benchmark datasets, i.e., OTB100 and UAV123, demonstrate that our TaiS consistently achieves high performance and updates the state of the arts in visual tracking. The source code of our approach will be made publicly available. Xingchao Liu, Ce Li 0002, Hongren Wang 0001, Xiantong Zhen, Baochang Zhang 0001, Qixiang Ye |
WACV | 6 |
| 2019 | Hierarchical residual stochastic networks for time series recognition
Chunyu Xie, Ce Li 0002, Baochang Zhang 0001, Lili Pan 0003, Qixiang Ye, Wei Chen 0016 |
Inf. Sci. | 5 |
| 2019 | Min-Entropy Latent Model for Weakly Supervised Object DetectionabstractWeakly supervised object detection is a challenging task when provided with image category supervision but required to learn, at the same time, object locations and object detectors. The inconsistency between the weak supervision and learning objectives introduces significant randomness to object locations and ambiguity to detectors. In this paper, a min-entropy latent model (MELM) is proposed for weakly supervised object detection. Min-entropy serves as a model to learn object locations and a metric to measure the randomness of object localization during learning. It aims to principally reduce the variance of learned instances and alleviate the ambiguity of detectors. MELM is decomposed into three components including proposal clique partition, object clique discovery, and object localization. MELM is optimized with a recurrent learning algorithm, which leverages continuation optimization to solve the challenging non-convexity problem. Experiments demonstrate that MELM significantly improves the performance of weakly supervised object detection, weakly supervised object localization, and image classification, against the state-of-the-art approaches. Fang Wan 0001, Pengxu Wei, Zhenjun Han, Jianbin Jiao, Qixiang Ye |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2019 | High performance person re-identification via a boosting ranking ensemble
Zhaoju Li, Zhenjun Han, Junliang Xing, Qixiang Ye, Xuehui Yu, Jianbin Jiao |
Pattern Recognit. | 4 |
| 2019 | Deep contour and symmetry scored object proposal
Wei Ke 0003, Jie Chen 0001, Qixiang Ye |
Pattern Recognit. Lett. | 3 |
| 2019 | Deep Manifold Structure Transfer for Action RecognitionabstractWhile intrinsic data structure in subspace provides useful information for visual recognition, it has not yet been well studied in deep feature learning for action recognition. In this paper, we introduce a new spatio-temporal manifold network (STMN) that leverages data manifold structures to regularize deep action feature learning, aiming at simultaneously minimizing the intra-class variations of learned deep features and alleviating the over-fitting problem. To this end, the manifold prior is imposed from the top layer of a convolutional neural network (CNN), and is propagated across convolutional layers during forward-backward propagation. The observed correspondence of manifold structures in the data space and feature space validates that the manifold priori can be transferred across CNN layers. STMN theoretically recasts the problem of transferring the data structure prior into the deep learning architectures as a projection over the manifold via an embedding method, which can be easily solved by an Alternating Direction Method of Multipliers and Backward Propagation (ADMM-BP) algorithm. STMN is generic in the sense that it can be plugged into various backbone architectures to learn more discriminative representation for action recognition. Extensive experimental results show that our method achieves comparable or even better performance as compared with the state-of-the-art approaches on four benchmark datasets. Ce Li 0002, Baochang Zhang 0001, Chen Chen 0001, Qixiang Ye, Jungong Han, Guodong Guo, Rongrong Ji |
IEEE Trans. Image Process. | 4 |
| 2018 | Image-Image Domain Adaptation With Preserved Self-Similarity and Domain-Dissimilarity for Person Re-IdentificationabstractPerson re-identification (re-ID) models trained on one domain often fail to generalize well to another. In our attempt, we present a "learning via translation" framework. In the baseline, we translate the labeled images from source to target domain in an unsupervised manner. We then train re-ID models with the translated images by supervised methods. Yet, being an essential part of this framework, unsupervised image-image translation suffers from the information loss of source-domain labels during translation. Our motivation is two-fold. First, for each image, the discriminative cues contained in its ID label should be maintained after translation. Second, given the fact that two domains have entirely different persons, a translated image should be dissimilar to any of the target IDs. To this end, we propose to preserve two types of unsupervised similarities, 1) self-similarity of an image before and after translation, and 2) domain-dissimilarity of a translated source image and a target image. Both constraints are implemented in the similarity preserving generative adversarial network (SPGAN) which consists of an Siamese network and a CycleGAN. Through domain adaptation experiment, we show that images generated by SPGAN are more suitable for domain adaptation and yield consistent and competitive re-ID accuracy on two large-scale datasets. Weijian Deng, Liang Zheng 0001, Qixiang Ye, Guoliang Kang, Yi Yang 0001, Jianbin Jiao |
CVPR | 3 |
| 2018 | Min-Entropy Latent Model for Weakly Supervised Object DetectionabstractWeakly supervised object detection is a challenging task when provided with image category supervision but required to learn, at the same time, object locations and object detectors. The inconsistency between the weak supervision and learning objectives introduces randomness to object locations and ambiguity to detectors. In this paper, a min-entropy latent model (MELM) is proposed for weakly supervised object detection. Min-entropy is used as a metric to measure the randomness of object localization during learning, as well as serving as a model to learn object locations. It aims to principally reduce the variance of positive instances and alleviate the ambiguity of detectors. MELM is deployed as two sub-models, which respectively discovers and localizes objects by minimizing the global and local entropy. MELM is unified with feature learning and optimized with a recurrent learning algorithm, which progressively transfers the weak supervision to object locations. Experiments demonstrate that MELM significantly improves the performance of weakly supervised detection, weakly supervised localization, and image classification, against the state-of-the-art approaches. Fang Wan 0001, Pengxu Wei, Jianbin Jiao, Zhenjun Han, Qixiang Ye |
CVPR | 5 |
| 2018 | Weakly Supervised Instance Segmentation Using Class Peak ResponseabstractWeakly supervised instance segmentation with image-level labels, instead of expensive pixel-level masks, remains unexplored. In this paper, we tackle this challenging problem by exploiting class peak responses to enable a classification network for instance mask extraction. With image labels supervision only, CNN classifiers in a fully convolutional manner can produce class response maps, which specify classification confidence at each image location. We observed that local maximums, i.e., peaks, in a class response map typically correspond to strong visual cues residing inside each instance. Motivated by this, we first design a process to stimulate peaks to emerge from a class response map. The emerged peaks are then back-propagated and effectively mapped to highly informative regions of each object instance, such as instance boundaries. We refer to the above maps generated from class peak responses as Peak Response Maps (PRMs). PRMs provide a fine-detailed instance-level representation, which allows instance masks to be extracted even with some off-the-shelf methods. To the best of our knowledge, we for the first time report results for the challenging image-level supervised instance segmentation task. Extensive experiments show that our method also boosts weakly supervised pointwise localization as well as semantic segmentation performance, and reports state-of-the-art results on popular benchmarks, including PASCAL VOC 2012 and MS COCO. Yanzhao Zhou, Yi Zhu 0004, Qixiang Ye, Qiang Qiu 0001, Jianbin Jiao |
CVPR | 3 |
| 2018 | Linear Span Network for Object Skeleton Detection
Chang Liu 0042, Wei Ke 0003, Qixiang Ye |
ECCV (2) | 4 |
| 2018 | Manifold constraint transfer for visual structure-driven optimization
Baochang Zhang 0001, Alessandro Perina, Ce Li 0002, Qixiang Ye, Vittorio Murino, Alessio Del Bue |
Pattern Recognit. | 4 |
| 2017 | SRN: Side-Output Residual Network for Object Symmetry Detection in the WildabstractIn this paper, we establish a baseline for object symmetry detection in complex backgrounds by presenting a new benchmark and an end-to-end deep learning approach, opening up a promising direction for symmetry detection in the wild. The new benchmark, named Sym-PASCAL, spans challenges including object diversity, multi-objects, part-invisibility, and various complex backgrounds that are far beyond those in existing datasets. The proposed symmetry detection approach, named Side-output Residual Network (SRN), leverages output Residual Units (RUs) to fit the errors between the object symmetry ground-truth and the outputs of RUs. By stacking RUs in a deep-to-shallow manner, SRN exploits the flow of errors among multiple scales to ease the problems of fitting complex outputs with limited layers, suppressing the complex backgrounds, and effectively matching object symmetry of different scales. Experimental results validate both the benchmark and its challenging aspects related to real-world images, and the state-of-the-art performance of our symmetry detection approach. The benchmark and the code for SRN are publicly available at https://github.com/KevinKecc/SRN. Wei Ke 0003, Jie Chen 0001, Jianbin Jiao, Guoying Zhao 0001, Qixiang Ye |
CVPR | 5 |
| 2017 | Self-Learning Scene-Specific Pedestrian Detectors Using a Progressive Latent ModelabstractIn this paper, a self-learning approach is proposed towards solving scene-specific pedestrian detection problem without any human annotation involved. The self-learning approach is deployed as progressive steps of object discovery, object enforcement, and label propagation. In the learning procedure, object locations in each frame are treated as latent variables that are solved with a progressive latent model (PLM). Compared with conventional latent models, the proposed PLM incorporates a spatial regularization term to reduce ambiguities in object proposals and to enforce object localization, and also a graph-based label propagation to discover harder instances in adjacent frames. With the difference of convex (DC) objective functions, PLM can be efficiently optimized with a concave-convex programming and thus guaranteeing the stability of self-learning. Extensive experiments demonstrate that even without annotation the proposed self-learning approach outperforms weakly supervised learning approaches, while achieving comparable performance with transfer learning and fully supervised approaches. Qixiang Ye, Tianliang Zhang 0003, Wei Ke 0003, Qiang Qiu 0001, Jie Chen 0001, Guillermo Sapiro, Baochang Zhang 0001 |
CVPR | 1 |
| 2017 | Oriented Response NetworksabstractDeep Convolution Neural Networks (DCNNs) are capable of learning unprecedentedly effective image representations. However, their ability in handling significant local and global image rotations remains limited. In this paper, we propose Active Rotating Filters (ARFs) that actively rotate during convolution and produce feature maps with location and orientation explicitly encoded. An ARF acts as a virtual filter bank containing the filter itself and its multiple unmaterialised rotated versions. During back-propagation, an ARF is collectively updated using errors from all its rotated versions. DCNNs using ARFs, referred to as Oriented Response Networks (ORNs), can produce within-class rotation-invariant deep features while maintaining inter-class discrimination for classification tasks. The oriented response produced by ORNs can also be used for image and object orientation estimation tasks. Over multiple state-of-the-art DCNN architectures, such as VGG, ResNet, and STN, we consistently observe that replacing regular filters with the proposed ARFs leads to significant reduction in the number of network parameters and improvement in classification performance. We report the best results on several commonly used benchmarks. Yanzhao Zhou, Qixiang Ye, Qiang Qiu 0001, Jianbin Jiao |
CVPR | 2 |
| 2017 | A scalable convolutional neural network for task-specified scenarios via knowledge distillationabstractIn this paper, we explore the redundancy in convolutional neural network, which scales with the complexity of vision tasks. Considering that many front-end visual systems are interested in only a limited range of visual targets, the removing of task-specified network redundancy can promote a wide range of potential applications. We propose a task-specified knowledge distillation algorithm to derive a simplified model with pre-set computation cost and minimized accuracy loss, which suits the resource constraint front-end systems well. Experiments on the MNIST and CIFAR10 datasets demonstrate the feasibility of the proposed approach as well as the existence of task-specified redundancy. Qixiang Ye, Zhenjun Han, Jianbin Jiao |
ICASSP | 3 |
| 2017 | Soft Proposal Networks for Weakly Supervised Object LocalizationabstractWeakly supervised object localization remains challenging, where only image labels instead of bounding boxes are available during training. Object proposal is an effective component in localization, but often computationally expensive and incapable of joint optimization with some of the remaining modules. In this paper, to the best of our knowledge, we for the first time integrate weakly supervised object proposal into convolutional neural networks (CNNs) in an end-to-end learning manner. We design a network component, Soft Proposal (SP), to be plugged into any standard convolutional architecture to introduce the nearly cost-free object proposal, orders of magnitude faster than state-of-the-art methods. In the SP-augmented CNNs, referred to as Soft Proposal Networks (SPNs), iteratively evolved object proposals are generated based on the deep feature maps then projected back, and further jointly optimized with network parameters, with image-level supervision only. Through the unified learning process, SPNs learn better object-centric filters, discover more discriminative visual evidence, and suppress background interference, significantly boosting both weakly supervised object localization and classification performance. We report the best results on popular benchmarks, including PASCAL VOC, MS COCO, and ImageNet. Yi Zhu 0004, Yanzhao Zhou, Qixiang Ye, Qiang Qiu 0001, Jianbin Jiao |
ICCV | 3 |
| 2017 | Keyword-driven image captioning via Context-dependent Bilateral LSTMabstractImage captioning has recently received much attention. Existing approaches, however, are limited to describing images with simple contextual information, which typically generate one sentence to describe each image with only a single contextual emphasis. In this paper, we address this limitation from a user perspective with a novel approach. Given some keywords as additional inputs, the proposed method would generate various descriptions according to the provided guidance. Hence, descriptions with different focuses can be generated for the same image. Our method is based on a new Context-dependent Bilateral Long Short-Term Memory (CDB-LSTM) model to predict a keyword-driven sentence by considering the word dependence. The word dependence is explored externally with a bilateral pipeline, and internally with a unified and joint training process. Experiments on the MS COCO dataset demonstrate that the proposed approach not only significantly outperforms the baseline method but also shows good adaptation and consistency with various keywords. Xiaodan Zhang 0003, Shengfeng He, Xinhang Song, Pengxu Wei, Shuqiang Jiang, Qixiang Ye, Jianbin Jiao, Rynson W. H. Lau |
ICME | 6 |
| 2017 | Beyond Group: Multiple Person Tracking via Minimal Topology-Energy-VariationabstractTracking multiple persons is a challenging task when persons move in groups and occlude each other. Existing group-based methods have extensively investigated how to make group division more accurately in a tracking-by-detection framework; however, few of them quantify the group dynamics from the perspective of targets' spatial topology or consider the group in a dynamic view. Inspired by the sociological properties of pedestrians, we propose a novel socio-topology model with a topology-energy function to factor the group dynamics of moving persons and groups. In this model, minimizing the topology-energy-variance in a two-level energy form is expected to produce smooth topology transitions, stable group tracking, and accurate target association. To search for the strong minimum in energy variation, we design the discrete group-tracklet jump moves embedded in the gradient descent method, which ensures that the moves reduce the energy variation of group and trajectory alternately in the varying topology dimension. Experimental results on both RGB and RGB-D data sets show the superiority of our proposed model for multiple person tracking in crowd scenes. Shan Gao 0003, Qixiang Ye, Junliang Xing, Arjan Kuijper, Zhenjun Han, Jianbin Jiao, Xiangyang Ji |
IEEE Trans. Image Process. | 2 |
| 2017 | Correlated Topic Vector for Scene ClassificationabstractScene images usually involve semantic correlations, particularly when considering large-scale image data sets. This paper proposes a novel generative image representation, correlated topic vector, to model such semantic correlations. Oriented from the correlated topic model, correlated topic vector intends to naturally utilize the correlations among topics, which are seldom considered in the conventional feature encoding, e.g., Fisher vector, but do exist in scene images. It is expected that the involvement of correlations can increase the discriminative capability of the learned generative model and consequently improve the recognition accuracy. Incorporated with the Fisher kernel method, correlated topic vector inherits the advantages of Fisher vector. The contributions to the topics of visual words have been further employed by incorporating the Fisher kernel framework to indicate the differences among scenes. Combined with the deep convolutional neural network (CNN) features and Gibbs sampling solution, correlated topic vector shows great potential when processing large-scale and complex scene image data sets. Experiments on two scene image data sets demonstrate that correlated topic vector improves significantly the deep CNN features, and outperforms existing Fisher kernel-based features. Pengxu Wei, Fang Wan 0001, Yi Zhu 0004, Jianbin Jiao, Qixiang Ye |
IEEE Trans. Image Process. | 6 |
| 2017 | Output Constraint Transfer for Kernelized Correlation Filter in TrackingabstractThe kernelized correlation filter (KCF) is one of the state-of-the-art object trackers. However, it does not reasonably model the distribution of correlation response during tracking process, which might cause the drifting problem, especially when targets undergo significant appearance changes due to occlusion, camera shaking, and/or deformation. In this paper, we propose an output constraint transfer (OCT) method that by modeling the distribution of correlation response in a Bayesian optimization framework is able to mitigate the drifting problem. OCT builds upon the reasonable assumption that the correlation response to the target image follows a Gaussian distribution, which we exploit to select training samples and reduce model uncertainty. OCT is rooted in a new theory which transfers data distribution to a constraint of the optimized variable, leading to an efficient framework to calculate correlation filters. Extensive experiments on a commonly used tracking benchmark show that the proposed method significantly improves KCF, and achieves better performance than other state-of-the-art trackers. To encourage further developments, the source code is made available. Baochang Zhang 0001, Xianbin Cao 0001, Qixiang Ye, Chen Chen 0001, LinLin Shen, Alessandro Perina, Rongrong Ji |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2016 | Person re-identification via adaboost ranking ensembleabstractMatching specific persons across scenes, known as person re-identification, is an important yet unsolved computer vision problem. Feature representation and metric learning are two fundamental factors in person re-identification. However, current person re-identification methods, which use single handcrafted feature with corresponding metric, could be not powerful enough when facing illumination, viewpoint and pose variations. Thus it inevitably produces suboptimal ranking lists. In this paper, we propose incorporating multiple features with metrics to build weak learners, and aggregate the base ranking lists by AdaBoost Ranking. Experiments on two commonly used datasets, VIPeR and CUHK01, show that our proposed approach greatly improves recognition rates over the state-of-the-art methods. Zhaoju Li, Zhenjun Han, Qixiang Ye |
ICIP | 3 |
| 2016 | Weakly supervised object detection with correlation and part suppressionabstractIn weakly supervised object detection, conventional methods treat object location in each image as a latent variable and use non-convex optimization to solve the latent variable. However, as the optimization objective is image-level instead of sample-level, the learning procedure tends to choose object parts as false positive samples. Furthermore, when multiple classes of objects appear in the same images, the models could invite class-correlations and lose discriminative capability. In this paper, we propose a simple but effective suppression strategy that mines hard negative samples in the learning procedure to ease the above problems. We propose using a spatial-voting strategy to help finding negative samples to suppress the impact of object parts. We also use regions from class-correlated images as negative samples to suppress the impact of class-correlations. Experiments show that our approach significantly improves the baseline by 6% and achieves state-of-the-art performance. Fang Wan 0001, Pengxu Wei, Zhenjun Han, Kun Fu 0001, Qixiang Ye |
ICIP | 5 |
| 2016 | Collective motion pattern inference via Locally Consistent Latent Dirichlet Allocation
Jialing Zou, Qixiang Ye, Yanting Cui, Fang Wan 0001, Kun Fu 0001, Jianbin Jiao |
Neurocomputing | 2 |
| 2015 | Pedestrian detection via PCA filters based convolutional channel featuresabstractIn this paper, we propose a kind of image representation, named PCA filters based convolutional channel features (PCA-CCF) for pedestrian detection. The motivation is to use the convolutional network architecture with orthogonal PCA filters to enhance the state-of-the-art aggregate channel features (ACF). In PCA-CCF, the convolutional operation improves the feature robustness to pedestrian local deformation. The learned PCA filters reduce the correlations among features of each channel, and therefore, improve feature discrimination capability. With the proposed PCA-CCF features and cascaded AdaBoost classifiers, we develop a coarse-to-fine pedestrian detection approach. Experiments show that such approach achieves 3.04%, 17.87% and 6.28% performance gain on the INRIA, Caltech Reasonable and Caltech Overall pedestrian datasets, respectively. Wei Ke 0003, Pengxu Wei, Qixiang Ye, Jianbin Jiao |
ICASSP | 4 |
| 2015 | Orientation robust object detection in aerial images using deep convolutional neural networkabstractDetecting objects in aerial images is challenged by variance of object colors, aspect ratios, cluttered backgrounds, and in particular, undetermined orientations. In this paper, we propose to use Deep Convolutional Neural Network (DCNN) features from combined layers to perform orientation robust aerial object detection. We explore the inherent characteristics of DC-NN as well as relate the extracted features to the principle of disentangling feature learning. An image segmentation based approach is used to localize ROIs of various aspect ratios, and ROIs are further classified into positives or negatives using an SVM classifier trained on DCNN features. With experiments on two datasets collected from Google Earth, we demonstrate that the proposed aerial object detection approach is simple but effective. Haigang Zhu, Weiqun Dai, Kun Fu 0001, Qixiang Ye, Jianbin Jiao |
ICIP | 5 |
| 2015 | Rich Image Description Based on RegionsabstractAbstract Automatically describing the content of an image is a fundamental problem in artificial intelligence that connects computer vision and natural language processing. In contrast to the previous image description methods that focus on describing the whole image, this paper presents a method of generating rich image descriptions from image regions. First, we detect regions with R-CNN (regions with convolutional neural network features) framework. We then utilize the RNN (recurrent neural networks) to generate sentences for image regions. Finally, we propose an optimization method to select one suitable region. The proposed model generates several sentence description of regions in an image, which has sufficient representative power of the whole image and contains more detailed information. Comparing to general image level description, generating more specific and accurate sentences on the different regions can satisfy more personal requirements for different people. Experimental evaluations validate the effectiveness of the proposed method. Xiaodan Zhang 0003, Xinhang Song, Xiong Lv, Shuqiang Jiang, Qixiang Ye, Jianbin Jiao |
ACM Multimedia | 5 |
| 2015 | Text Detection and Recognition in Imagery: A SurveyabstractThis paper analyzes, compares, and contrasts technical challenges, methods, and the performance of text detection and recognition research in color imagery. It summarizes the fundamental problems and enumerates factors that should be considered when addressing these problems. Existing techniques are categorized as either stepwise or integrated and sub-problems are highlighted including text localization, verification, segmentation and recognition. Special issues associated with the enhancement of degraded text and the processing of video text, multi-oriented, perspectively distorted and multilingual text are also addressed. The categories and sub-categories of text are illustrated, benchmark datasets are enumerated, and the performance of the most representative approaches is compared. This review provides a fundamental comparison and analysis of the remaining problems in the field. Qixiang Ye, David S. Doermann |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2015 | Real-Time Multipedestrian Tracking in Traffic Scenes via an RGB-D-Based Layered Graph ModelabstractMultipedestrian tracking in traffic scenes is challenging due to cluttered backgrounds and serious occlusions. In this paper, we propose a layered graph model in image (RGB) and depth (D) domains for real-time robust multipedestrian tracking. The motivation is to investigate high-level constraints in RGB-D data association and to improve the optimization from the trajectory level to the layer level. To construct a layered graph, we define constraints in the depth domain so that pedestrian objects in the image domain are assigned to proper layers. We use pedestrian detection responses in the RGB domain as graph nodes, and we integrate 3-D motion, appearance, and depth features as graph edges. An online updating depth factor is defined to describe the depth relationships among the observations in and out of the layers, and the occlusion issue is processed with an analytical layer-level strategy. With a heuristic label switching algorithm, multiple pedestrian objects are optimally associated and tracked. Experiments and comparison on five public data sets show that our proposed approach significantly reduces pedestrian's ID switch and improves tracking accuracy in the cases of serious occlusions. Shan Gao 0003, Zhenjun Han, Ce Li 0005, Qixiang Ye, Jianbin Jiao |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2014 | Robust scene text detection using integrated feature discriminationabstractScene text detection in images of cluttered backgrounds and/or multilingual context is very challenging. In this paper, we propose a discriminative approach that integrates appearance and consensus features for robust scene text detection. We propose an integrated discrimination model to perform text classification as well as control component grouping. We design shape, stroke and structural features to describe text component appearance and the consensus among them. Experimental results on three public datasets show that the proposed approach is robust to cluttered backgrounds, and is applicable in multilingual environments. Qixiang Ye, David S. Doermann |
ICIP | 1 |
| 2014 | A cluster specific latent dirichlet allocation model for trajectory clustering in crowded videosabstractTrajectory analysis in crowded video scenes is challenging as trajectories obtained by existing tracking algorithms are often fragmented. In this paper, we propose a new approach to do trajectory inference and clustering on fragmented trajectories, by exploring a cluster specific Latent Dirichlet Allocation(CLDA) model. LDA models are widely used to learn middle level trajectory features and perform trajectory inference. However, they often require scene priors in the learning or inference process. Our cluster specific LDA model addresses this issue by using manifold based clustering as initialization and iterative statistical inference as optimization. The output middle level features of CLDA are input to a clustering algorithm to obtain trajectory clusters. Experiments on a public dataset show the effectiveness of our approach. Jialing Zou, Yanting Cui, Fang Wan 0001, Qixiang Ye, Jianbin Jiao |
ICIP | 4 |
| 2014 | Locality-Constrained Sparse Reconstruction for Trajectory ClassificationabstractTrajectory classification has been extensively investigated in recent years, however, problems remain when processing incomplete trajectories of noises and local variations. In this paper, we propose a Locality-constrained Sparse Reconstruction (LSR) approach that explores both sparsity and local adaptability for robust trajectory classification. A trajectory dictionary with locality constrains is constructed with track lets partitioned from collected trajectories by control points of cubic B-spline curves. On the dictionary, the proposed LSR is used to calculate a discriminate code matrix. Then, a loss weighted decoding strategy is employed to perform multi-class trajectory classification. In addition, the approach can be used for anomalous trajectory detection with a thresholding strategy. Experiments on two datasets show that the results of the LSR approach improve the state of the art. Ce Li 0005, Zhenjun Han, Qixiang Ye, Shan Gao 0003, Lijin Pang, Jianbin Jiao |
ICPR | 3 |
| 2014 | A Belief Based Correlated Topic Model for Trajectory Clustering in Crowded Video ScenesabstractTrajectory clustering in crowded video scenes is very challenging. In this paper, we propose to use a belief based correlated topic model (BCTM) to learn discriminative middle level features for trajectory clustering. By constructing a scene prior based joint Gaussian distribution, the BCTM can uncover relations between trajectory clusters and the middle level features using a parameter estimation procedure. The method has distinct advantages over Correlated Topic Model (CTM) and Random Field Topic (RFT) model previously proposed. The inputs to the BCTM are either full trajectories or trajectory fragments obtained with an existing tracking algorithm. The output BCTM features are input to a hierarchical clustering algorithm to obtain trajectory clusters. Experiments on three benchmark datasets show that the proposed BCTM and trajectory clustering approach improves the state of the art. Jialing Zou, Qixiang Ye, Yanting Cui, David S. Doermann, Jianbin Jiao |
ICPR | 2 |
| 2013 | Robust Visual Object Tracking via Sparse Representation and Reconstruction
Zhenjun Han, Qixiang Ye, Jianbin Jiao |
CAIP (2) | 2 |
| 2013 | Visual abnormal behavior detection based on trajectory sparse reconstruction analysis
Ce Li 0005, Zhenjun Han, Qixiang Ye, Jianbin Jiao |
Neurocomputing | 3 |
| 2013 | Human Detection in Images via Piecewise Linear Support Vector MachinesabstractHuman detection in images is challenged by the view and posture variation problem. In this paper, we propose a piecewise linear support vector machine (PL-SVM) method to tackle this problem. The motivation is to exploit the piecewise discriminative function to construct a nonlinear classification boundary that can discriminate multiview and multiposture human bodies from the backgrounds in a high-dimensional feature space. A PL-SVM training is designed as an iterative procedure of feature space division and linear SVM training, aiming at the margin maximization of local linear SVMs. Each piecewise SVM model is responsible for a subspace, corresponding to a human cluster of a special view or posture. In the PL-SVM, a cascaded detector is proposed with block orientation features and a histogram of oriented gradient features. Extensive experiments show that compared with several recent SVM methods, our method reaches the state of the art in both detection accuracy and computational efficiency, and it performs best when dealing with low-resolution human regions in clutter backgrounds. Qixiang Ye, Zhenjun Han, Jianbin Jiao, Jianzhuang Liu |
IEEE Trans. Image Process. | 1 |
| 2012 | Pedestrian detection via part-based topology modelabstractIn this paper, we propose a part-based topology model and a pedestrian detection method, which obviously improve the detection accuracy. In Our method, pedestrian is divided into several parts. Firstly, histogram of oriented gradients (HOG) features and linear support vector machine (SVM) classifier are used to detect pedestrian parts. Secondly, a novel binary descriptor called log-polar pattern (LPP) is proposed to represent the spatial relation of a part pair. Then multiple LPPs are combined as a log-polar topology pattern (LTP) to model the global topology of a pedestrian. Finally, we put the LTP into One-Class SVM (OC-SVM) to determine whether the detected parts indicate a pedestrian or not. Experiments in INRIA dataset show that our method is robust to occlusion and multi-postures, which obviously reduces the miss rate. Wen Gao 0001, Qixiang Ye, Jianbin Jiao |
ICIP | 3 |
| 2012 | Evaluation of local feature descriptors and their combination for pedestrian representation
Jixiang Liang, Qixiang Ye, Jie Chen 0001, Jianbin Jiao |
ICPR | 2 |
| 2012 | Pedestrian detection in images via cascaded L1-norm minimization learning method
Jianbin Jiao, Baochang Zhang 0001, Qixiang Ye |
Pattern Recognit. | 4 |
| 2012 | Pedestrian Detection in Video Images via Error Correcting Output Code Classification of Manifold SubclassesabstractPedestrian detection in images and video frames is challenged by the view and posture problem. In this paper, we propose a new pedestrian detection approach by error correcting output code (ECOC) classification of manifold subclasses. The motivation is that pedestrians across views and postures form a manifold and that the ECOC method constructs a nonlinear classification boundary that can discriminate the manifold from negative samples. The pedestrian manifold is first constructed with a local linear embedding algorithm and then divided into subclasses with a -means clustering algorithm. The neighboring relationships of these subclasses are used to make the encoding rule for ECOCs, which we use to train multiple base classifiers with histogram of oriented gradient features and linear support vector machines. In the detection procedure, image windows are tested with all base classifiers, and their output codes are fed into an ECOC decoding procedure to decide whether it is a pedestrian or not. Experiments on three data sets show that the results of our approach improve the state of the art. Qixiang Ye, Jixiang Liang, Jianbin Jiao |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2011 | Abnormal Behavior Detection via Sparse Reconstruction Analysis of TrajectoryabstractThis paper proposes a new method for abnormal behavior detection in surveillance videos via sparse reconstruction analysis. The motion trajectories of objects are firstly defined as fixed-length parametric vectors based on approximating cubic B-spline curves. Then the vectors are classified as behavior patterns and finally distinguished between normal and abnormal behaviors based on sparse reconstruction analysis, in which a classifier is constructed with sparse linear reconstruction coefficients by computing L1-norm minimization and sparse reconstruction residuals learning from labeled training samples. Experimental results on public dataset show the effectiveness of the proposed approach. Ce Li 0005, Zhenjun Han, Qixiang Ye, Jianbin Jiao |
ICIG | 3 |
| 2011 | Fast Pedestrian Detection with Laser and Image Data FusionabstractIn this paper, we proposed a pedestrian detection system based on laser and image data fusion. The high speed of laser data based location and precise of image based classification are fully explored. First, laser scanner point data is clustered into segments, each of which implies a pedestrian candidate. Then, the segments are projected to the image domain to form regions of interest (ROI) on the image, given camera calibration parameters. Finally two SVM classifiers on Histogram of Oriented Gradient (HOG) features are used to precisely locate pedestrians on the ROI. Experiments report over 30 times higher speed than the state-of-the-art method and a comparable detection rate. Jixiang Liang, Qixiang Ye, Zhenjun Han, Jianbin Jiao |
ICIG | 3 |
| 2011 | A fast object tracking approach based on sparse representationabstractThis paper proposes a new approach based on object sparse representation (OSR) for object tracking. The OSR method implemented by L1-norm minimization is robust to the partial occlusion and deterioration in object images. Firstly, we dynamically construct a set of samples in a predicted searching window in a new video frame, on which the sparse representation of the tracked object can be calculated by the OSR method. This procedure can automatically select the subset of the samples as a basis which most compactly expresses the object with small residuals and rejects all other possible but less compact representations. In terms of this sparse and compact representation, the instantaneous tracking result is achieved in the new video frame. Extensive comparative experiments demonstrate the effectiveness of the proposed approach especially in occlusion context. Zhenjun Han, Jianbin Jiao, Qixiang Ye |
ICIP | 3 |
| 2011 | Nonlinear L1-norm minimization learning for human detectionabstractView, appearance and pose variations make it difficult to detect human objects only by using linear classification methods. Inspired by the successful applications of L1-norm minimization learning (LML) for human detection, we propose a new nonlinear L1-norm minimization learning method (NL-LML). It integrates a nonlinear transformation with an LML optimization model for human detection. The NL-LML method first maps the samples into a space based on the kernel function, and then combines the reformulated samples in the transformed space with the LML model to learn a classifier. Histograms of orientated gradient (HOG) features are used as the feature descriptors, and the sliding window scheme is adopted to detect humans in images. Experiments on two human datasets validate the efficiency and effectiveness of the proposed method. Jianbin Jiao, Qixiang Ye |
ICIP | 3 |
| 2011 | Combined feature evaluation for adaptive visual object tracking
Zhenjun Han, Qixiang Ye, Jianbin Jiao |
Comput. Vis. Image Underst. | 2 |
| 2011 | Visual object tracking via sample-based Adaptive Sparse Representation (AdaSR)
Zhenjun Han, Jianbin Jiao, Baochang Zhang 0001, Qixiang Ye, Jianzhuang Liu |
Pattern Recognit. | 4 |
| 2010 | Cascaded L1-norm Minimization Learning (CLML) classifier for human detectionabstractThis paper proposes a new learning method, which integrates feature selection with classifier construction for human detection via solving three optimization models. Firstly, the method trains a series of weak-classifiers by the proposed L1-norm Minimization Learning (LML) and min-max penalty function models. Secondly, the proposed method selects the weak-classifiers by using the integer optimization model to construct a strong classifier. The L1-norm minimization and integer optimization models aim to find the minimal VC-dimension for weak and strong classifiers respectively. Finally, the method constructs a cascade of LML (CLML) classifier to reach higher detection rates and efficiency. Histograms of Oriented Gradients features of variable-size blocks (v-HOG) are employed as human representation to verify the proposed method. Experiments conducted on INRIA human test set show more superior detection rates and speed than state-of-the-art methods. Baochang Zhang 0001, Qixiang Ye, Jianbin Jiao |
CVPR | 3 |
| 2010 | Human detection in images via L1-norm Minimization LearningabstractIn recent years, sparse representation originating from signal compressed sensing theory has attracted increasing interest in computer vision research community. However, to our best knowledge, no previous work utilizes L1-norm minimization for human detection. In this paper we develop a novel human detection system based on L1-norm Minimization Learning (LML) method. The method is on the observation that a human object can be represented by a few features from a large feature set (sparse representation). And the sparse representation can be learned from the training samples by exploiting the L1-norm Minimization principle, which can also be called feature selection procedure. This procedure enables the feature representation more concise and more adaptive to object occlusion and deformation. After that a classifier is constructed by linearly weighting features and comparing the result with a calculated threshold. Experiments on two datasets validate the effectiveness and efficiency of the proposed method. Baochang Zhang 0001, Qixiang Ye, Jianbin Jiao |
ICASSP | 3 |
| 2010 | Fast pedestrian detection with multi-scale orientation features and two-stage classifiersabstractIn this paper, we propose an approach for fast pedestrian detection in images. Inspired by the histogram of oriented gradient (HOG) features, a set of multi-scale orientation (MSO) features are proposed as the feature representation. The features are extracted on square image blocks of various sizes (called units), containing coarse and fine features in which coarse ones are the unit orientations and fine ones are the pixel orientation histograms of the unit. A cascade of Adaboost is employed to train classifiers on the coarse features, aiming to high detection speed. A greedy searching algorithm is employed to select fine features, which are input into SVMs to train the fine classifiers, aiming to high detection accuracy. Experiments report that our approach obtains state-of-art results with 12.4 times faster than the SVM+HOG method. Qixiang Ye, Jianbin Jiao, Baochang Zhang 0001 |
ICIP | 1 |
| 2009 | A framework for flexible summarization of racquet sports video using multiple modalities
Chunxi Liu, Qingming Huang, Shuqiang Jiang, Liyuan Xing, Qixiang Ye, Wen Gao 0001 |
Comput. Vis. Image Underst. | 5 |
| 2009 | A configurable method for multi-style license plate recognition
Jianbin Jiao, Qixiang Ye, Qingming Huang |
Pattern Recognit. | 2 |
| 2008 | Online feature evaluation for object tracking using Kalman FilterabstractAn online feature evaluation method for visual object tracking is put forward in this paper. Firstly, a combined feature set is built using color histogram (HC) bins and gradient orientation histogram (HOG) bins considering the color and contour representation of an object respectively. Then a novel method is proposed to evaluate the features’ weights in a tracking process using Kalman Filter, which is used to comprise the inter-frame predication and single-frame measurement of features’ discriminative power. In this way, we extend the traditional filter framework from modeling motion states to modeling feature evaluation. Experiments show this method can greatly improve the tracking stabilization when objects go across complex backgrounds. Zhenjun Han, Qixiang Ye, Jianbin Jiao |
ICPR | 2 |
| 2007 | Multi-posture Human Detection in Video Frames by Motion Contour Matching
Qixiang Ye, Jianbin Jiao |
ACCV (1) | 1 |
| 2007 | Text detection and restoration in natural scene images
Qixiang Ye, Jianbin Jiao |
J. Vis. Commun. Image Represent. | 1 |
| 2006 | An effective method to detect and categorize digitized traditional Chinese paintings
Shuqiang Jiang, Qingming Huang, Qixiang Ye, Wen Gao 0001 |
Pattern Recognit. Lett. | 3 |
| 2005 | Playfield Detection Using Adaptive GMM and Its ApplicationabstractPlayfield detection is a key step in sports video content analysis, since many semantic clues could be inferred from it. In this paper we propose an adaptive GMM based algorithm for playfield detection. Its advantages are twofold. First, it can update model parameters by the incremental expectation maximization (IEM) algorithm, which enables the model to adapt to the playfield variation with time; Second, online training is performed, which saves buffer for training samples. Then, the playfield detection results are applied in recognizing the key zone of the current playfield in soccer video, in which a fast algorithm based on playfield contour and least square is proposed. Experimental results show that the proposed algorithms are encouraging. Yang Liu 0006, Shuqiang Jiang, Qixiang Ye, Wen Gao 0001, Qingming Huang |
ICASSP (2) | 3 |
| 2005 | Exciting event detection in broadcast soccer video with mid-level description and incremental learningabstractIn this paper, we propose a method for exciting event detection in broadcast soccer video with mid-level description and SVM-based incremental learning. In the method, video frames are firstly classified and grouped into views in terms of low-level playfield features. Mid-level description including view label, motion descriptor and shot descriptor are then extracted to present the characteristics of a view. By using the fixed temporal structure of views, SVM classification models are constructed to detected exciting events in a soccer match. In the view classification and event detection procedures, SVM-based incremental learning method is explored to improve the extensibility of view classification and event detection. Experiments on real soccer video programs demonstrate encouraging results. Qixiang Ye, Qingming Huang, Wen Gao 0001, Shuqiang Jiang |
ACM Multimedia | 1 |
| 2005 | Fast and robust text detection in images and video frames
Qixiang Ye, Qingming Huang, Wen Gao 0001, Debin Zhao |
Image Vis. Comput. | 1 |
| 2004 | Automatic text segmentation from complex backgroundabstractIn this paper, we proposed an automatic method to segment text from complex background for recognition task. First, a rule-based sampling method is proposed to get portion of the text pixels. Then, the sampled pixels are used for training Gaussian mixture models of intensity and hue components in HSI color space. Finally, the trained GMMs together with the spatial connectivity information are used for segment all of text pixels form their background. We used the word recognition rate to evaluate the segmentation result. Experiments results show that the proposed algorithm can work fully automatically and performs much better than the traditional methods. Qixiang Ye, Wen Gao 0001, Qingming Huang |
ICIP | 1 |
| 2004 | A new method to segment playfield and its applications in match analysis in sports videoabstractWith the growing popularity of digitized sports video, automatic analysis of them need be processed to facilitate semantic summarization and retrieval. Playfield plays the fundamental role in automatically analyzing many sports programs. Many semantic clues could be inferred from the results of playfield segmentation. In this paper, a novel playfield segmentation method based on Gaussian mixture models (GMMs) is proposed. Firstly, training pixels are automatically sampled from frames. Then, by supposing that field pixels are the dominant components in most of the video frames, we build the GMMs of the field pixels and use these models to detect playfield pixels. Finally region-growing operation is employed to segment the playfield regions from the background. Experimental results show that the proposed method is robust to various sports videos even for very poor grass field conditions. Based on the results of playfield segmentation, match situation analysis is investigated, which is also desired for sports professionals and longtime fanners. The results are encouraging. Shuqiang Jiang, Qixiang Ye, Wen Gao 0001, Tiejun Huang 0001 |
ACM Multimedia | 2 |
| 2003 | Color image segmentation using density-based clusteringabstractColor image segmentation is an important but still open problem in image processing. We propose a method for this problem by integrating the spatial connectivity and color features of the pixels. Considering that an image can be regarded as a dataset in which each pixel has a spatial location and a color value, color image segmentation can be obtained by clustering these pixels into different groups of coherent spatial connectivity and color. To discover the spatial connectivity of the pixels, density-based clustering is employed, which is an effective clustering method used in data mining for discovering spatial databases. The color similarity of the pixels is measured in Munsell (HVC) color space whose perceptual uniformity ensures the color change in the segmented regions is smooth in terms of human perception. Experimental results using the proposed method demonstrate encouraging performance. Qixiang Ye, Wen Gao 0001, Wei Zeng 0006 |
ICASSP (3) | 1 |
| 2003 | Color image segmentation using density-based clusteringabstractColor image segmentation is an important but still open problem in image processing. In this paper, we propose a method for this problem by integrating the spatial connectivity and color feature of the pixels. Considering that an image can be regarded as a dataset in which each pixel has a spatial location and a color value, color image segmentation can be obtained by clustering these pixels into different groups of coherent spatial connectivity and color. To discover the spatial connectivity of the pixels, density-based clustering is employed, which is an effective clustering method used in data mining for discovering spatial databases. Color similarity of the pixels is measured in Munsell (HVC) color space whose perceptual uniformity ensures the color change in the segmented regions is smooth in terms of human perception. Experimental results using proposed method demonstrate encouraging performance. Qixiang Ye, Wen Gao 0001, Wei Zeng 0006 |
ICME | 1 |
| 2003 | Objectionable Image Recognition System in Compression Domain
Qixiang Ye, Wen Gao 0001, Wei Zeng 0006, Weiqiang Wang 0001, Yang Liu 0006 |
IDEAL | 1 |