VLDB 2026 Research / reviewers in the wild / expert
Heqian Qiu
dblp:234/1711
· DBLP profile ↗
60ranked-venue papers
10as first author
55since 2021 · last 2026
0000-0002-0963-0311ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 47 · 9 first-author · 42 since 2021Artificial intelligence and machine learning · 18 · 3 first-author · 17 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Parameter Merging with Gradient-Guided Supermasks in Online Continual LearningabstractOnline continual learning (OCL) aims at learning a non-stationary data stream in a way of reading each data sample only once, and hence suffers from the trade-off of catastrophic forgetting and insufficient learning. In this work, we firstly analytically establish relationship between loss functions and model parameters from the Bayesian perspective. Based on our analysis, we subsequently propose a parameter merging method with gradient-guided supermasks. Our method leverages 1-order and 2-order gradient information to construct supermasks that determine the merging weights between the old and new models. Our method performs direct arithmetic operations on parameters to update models, beyond traditional gradient descent. We further discover that a widely-used premise that 1-order gradients can be negligible is invalid in OCL, due to slow convergence incurred by insufficient learning. Additionally, we utilize a dual-model dual-view distillation strategy that can align output distributions of the new and merged models for each sample, further enhancing model performance. Extensive experiments are conducted on four benchmarks in OCL settings, including CIFAR-10, CIFAR-100, Tiny-ImageNet, and ImageNet-100. Experimental results demonstrate that our method is effective, and achieves a substantial boost over previous methods. Benliu Qiu, Heqian Qiu, Lanxiao Wang, Taijin Zhao, Lili Pan 0001, Hongliang Li 0001 |
AAAI | 2 |
| 2026 | Ego-PMOVE: Prompt-aware Mixture of View Experts Network for Egocentric Gaze PredictionabstractEgocentric gaze prediction serves as a critical indicator for decoding human visual attention and cognitive processes, but its inherently limited field of view creates prediction challenges. Although exo-view data provides supplementary contextual information, it exhibits significant spatial and semantic gaps. Existing methods focus solely on isolated feature encoding in single-view paradigms, neglecting cross-view gaze correlations. To make up for this gap, we make the first exploration of cross-view gaze relationship for egocentric gaze prediction, and propose Ego-PMOVE, a novel Prompt-aware Mixture of View Experts network. Unlike prior cross-view studies that forcibly align cross-view features thereby introducing inference noise, we leverage the popular Mixture-of-Experts (MoE) and a set of flexible prompts to disentangle features from different views into three parallel experts: a view-shared expert directly modeling common semantic relationships, a view-discrepancy expert adaptively adjusting the spatial position, scale and shifts based on different view-specific features, and an egocentric expert extracting independent features to compensate for the case of missing exocentric data. To balance these experts, we further design a soft router to dynamically weight them for mining useful information while suppressing noise. A view-query gaze decoder then generates view-specific gaze attention maps, jointly optimized by gaze-heamap and cross-view contrastive loss that regularize both shared and divergent features for accurate gaze prediction. Extensive experiments across the multi-view EgoMe dataset and single-view Ego4D and EGTEA Gaze++ datasets demonstrate the effectiveness and generalizability of our approach. Heqian Qiu, Lanxiao Wang, Taijin Zhao, Zhaofeng Shi, Linfeng Xu 0001, Hongliang Li 0001 |
AAAI | 1 |
| 2026 | Bridging the Gap between Vision and Text for unsupervised text-only captioning
Lanxiao Wang, Heqian Qiu, Haitao Wen, Fanman Meng, Qingbo Wu 0001, Hongliang Li 0001 |
Pattern Recognit. | 2 |
| 2026 | Zippo: RGB-Alpha Joint Modeling With a Unified Diffusion ModelabstractRecent advances in generative models have sparked growing interest in moving beyond pure image generation toward transparent image generation, i.e., joint generation of image and its alpha mask. However, most existing approaches adopt a two-stage pipeline, where a diffusion-based model first generates an RGB image and a subsequent matting head predicts the alpha mask. This separation not only leads to error accumulation and inaccurate predictions but also overlooks the intrinsic correlation between the cross-modal data. In this work, we introduce Zippo, a unified diffusion framework, zipping color and transparency distributions into a single diffusion model, by learning joint distribution of RGB image and alpha mask. Zippo not only generates high-fidelity images but also produces plausible and sharp alpha masks. In practice, Zippo inflates the latent space into a unified representation that encodes cross-modal data, and builds upon it with a modality-aware diffusion process that flexibly switches between RGB and alpha domains. In this process, conditioning on one modality while denoising the other allows the model to generate RGB images from alpha masks and predict transparency from input images. In addition to single-modality prediction, we further design a modality-aware noise reassignment strategy to empower Zippo with the joint generation capability of RGB images and their corresponding alpha masks under text guidance. With these techniques, Zippo supports a wide range of transparent image generation tasks, including image-alpha joint generation, image matting, and alpha mask conditioned image generation. Extensive experiments demonstrate that Zippo not only delivers superior visual fidelity but also achieves competitive performance in visual downstream prediction, highlighting joint image-alpha modeling as a powerful alternative to traditional paradigms. Kangyang Xie, Chenchen Jing, Cheng Peng 0011, Ming Yang 0007, Heqian Qiu, Hongliang Li 0001, Hao Chen 0041 |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2025 | Unsupervised Ego- and Exo-centric Dense Procedural Activity Captioning via Gaze Consensus AdaptationabstractEven from an early age, humans naturally adapt between exocentric (Exo) and egocentric (Ego) perspectives to understand daily procedural activities. Inspired by this cognitive ability, we propose a novel Unsupervised Ego-Exo Dense Procedural Activity Captioning (UE^2 DPAC) task, which aims to transfer knowledge from the labeled source view to predict the time segments and descriptions of action sequences for the target view without annotations. Despite previous works endeavoring to address the fully-supervised single-view or cross-view dense video captioning, they lapse in the proposed task due to the significant inter-view gap caused by temporal misalignment and irrelevant object interference. Hence, we propose a Gaze Consensus-guided Ego-Exo Adaptation Network (GCEAN) that injects the gaze information into the learned representations for the fine-grained Ego-Exo alignment. Specifically, we propose a Score-based Adversarial Learning Module (SALM) that incorporates a discriminative scoring network and compares the scores of distinct views to learn unified view-invariant representations from a global level. Then, the Gaze Consensus Construction Module (GCCM) utilizes the gaze to progressively calibrate the learned representations to highlight the regions of interest and extract the corresponding temporal contexts. Moreover, we adopt hierarchical gaze-guided consistency losses to construct gaze consensus for the explicit temporal and spatial adaptation between the source and target views. To support our research, we propose a new EgoMe-UE^2 DPAC benchmark, and extensive experiments demonstrate the effectiveness of our method, which outperforms many related methods by a large margin. Code is available at https://github.com/ZhaofengSHI/GCEAN. Zhaofeng Shi, Heqian Qiu, Lanxiao Wang, Qingbo Wu 0001, Fanman Meng, Hongliang Li 0001 |
ACM Multimedia | 2 |
| 2025 | DBAB: A Dual-Branch Adaptive Balance Framework with Optimized Plasticity Branch for Class-Incremental LearningabstractIn the context of class-incremental learning, the primary challenge for models is to overcome catastrophic forgetting. Leveraging the strong generalization ability of frozen pre-trained models can significantly enhance model performance and alleviate catastrophic forgetting during training. To enable models to better adapt to downstream tasks, fine-tuning pre-trained models for new tasks is a common approach. However, current works struggle to balance plasticity and generalization performance after fine-tuning, as fine-tuning causes model parameters to overwrite knowledge of old tasks. This paper proposes a Dual-Branch Adaptive Balance (DBAB) framework, which consists of a plasticity branch fine-tuned from a pre-trained model and optimized for downstream tasks, and a generalization branch with frozen pre-trained parameters. The framework designs an adaptive balance mechanism for the dual branches, introduces learnable balance coefficients to dynamically fuse class prototype distances from both branches, and devises a loss function for training and regularizing the balance coefficients. This ensures a better balance between plasticity and generalization during the incremental learning process. To optimize the plasticity branch in the DBAB framework, an Adaptive Plasticity Module (APM) is proposed. Considering the heterogeneity of embedding distributions in continuous learning of downstream tasks, APM employs Mahalanobis distance for anisotropic feature alignment, uses a covariance matrix to dynamically adapt to the heterogeneous distributions of new tasks, and improves and stabilizes the Mahalanobis distance-based classification method.Experimental results show that DBAB outperforms multiple state-of-the-art (SOTA) methods on benchmark datasets such as CIFAR100 and CUB200, demonstrating significant performance improvements. Heqian Qiu, Chenghao Qi, Ruisong Dai, Hongliang Li 0001 |
MMSP | 2 |
| 2025 | OrthCal: Synergizing Orthogonal Contrastive Learning and Prototype Calibration for Few-Shot Class-Incremental LearningabstractFew-Shot Class-Incremental Learning (FSCIL) requires models to progressively learn novel classes with limited samples while mitigating catastrophic forgetting of base classes. Existing methods face dual challenges: novel classes are prone to misclassification into base classes because the strong discriminability of base classes distracts the classification of novel classes, and the feature space lacks sufficient generalization capacity. This paper proposes the OrthCal framework built on the deep integration of orthogonal contrastive learning and a prototype calibration strategy to improve the performance during incremental sessions. Our three-stage optimization includes: 1) pretraining with hybrid supervised and self-supervised contrastive learning to construct geometrically constrained orthogonal pseudo-targets. 2) Dynamic prototype calibration, using semantic similarity among base classes to adjust novel class prototypes without additional training. 3) Hybrid loss design optimizing orthogonality constraints, perturbation-sensitive contrastive loss, and calibrated prototypes jointly to address challenges arising from data limitations during incremental sessions. Experiments on miniImageNet and CIFAR100 demonstrate that OrthCal achieves state-of-the-art performance. The framework provides a unified solution for feature space optimization and prototype calibration in FSCIL. Ruisong Dai, Chenghao Qi, Heqian Qiu, Hongliang Li 0001 |
MMSP | 5 |
| 2025 | D3Net: Dual-Path Decoupling-Distillation for Adaptive Fusion in Continual Egocentric LearningabstractEgocentric continual action recognition faces severe challenges such as sudden viewpoint changes, occlusions, and complex backgrounds. In such scenarios, relying solely on visual modalities is susceptible to interference and lacks sufficient recognition robustness. To overcome the limitations of unimodal approaches, multimodal fusion methods are widely adopted, significantly enhancing recognition performance. However, existing multimodal schemes generally suffer from insufficient exploration of cross-modal complementarity and the vulnerability of modal independence. To address this, this paper proposes a Dual-path Decoupling-Distillation NetWork (D3Net), aiming to achieve more effective dynamic fusion of modal information and knowledge transfer.D3Net first explicitly separates the shared and private features of modalities through a dual-path decoupling module, combined with a dynamic gating mechanism to adaptively adjust the modal fusion weights. Secondly, it designs a complementary distillation module, leveraging cross-modal contrastive learning to effectively mitigate the issues of poor unimodal robustness and vulnerability to interference. Finally, through a cross-task distillation mechanism, it efficiently extracts knowledge from old tasks, alleviating the catastrophic forgetting problem during learning. Experimental results demonstrate that D3Net achieves an average accuracy of 83.97% under the 8×4 task configuration on the UESTC MMEA CL dataset, surpassing baseline method by 5.17%. Chenghao Qi, Heqian Qiu, Zhaofeng Shi, Lanxiao Wang, Hongliang Li 0001 |
MMSP | 2 |
| 2025 | Efficient Polyp Detection via Wavelet-Driven Boundary Enhancement and Temporal ConsistencyabstractAccurate early detection of polyps plays a critical role in preventing, diagnosing, and treating colorectal cancer. Although significant progress has been made, the accurate and efficient detection of polyps remains a challenging task. Existing image-based methods, while computationally efficient, typically rely on single-frame inputs and struggle to handle polyps with ambiguous boundaries or varying sizes. On the other hand, video-based approaches leverage temporal information to improve detection robustness, but often incur high computational costs and compromise real-time performance due to the need to process multiple frames simultaneously. Moreover, both types of methods are susceptible to dynamic artifacts caused by endoscopic camera movement, which can lead to polyp-like false positives. To address these issues, we propose BEC-Net, a novel Boundary-Enhanced network with adjacent-Frame Contrastive Learning for accurate and efficient polyp detection. Specifically, we design a Wavelet-Based Boundary-Aware feature Fusion (WBAF) module to enhance the representation of polyp boundaries and improve generalization across diverse appearances. To accommodate scale variation, we introduce a Context-Aware Gated Aggregation (CAGA) module that adaptively integrates multi-scale contextual information. Furthermore, we propose an Adjacent-Frame Contrastive Learning (AFCL) strategy that utilizes temporal consistency between adjacent frames to suppress polyp-like artifacts without increasing inference cost. Extensive experiments on large-scale colonoscopy benchmarks demonstrate that our method outperforms state-of-the-art approaches in both accuracy and real-time performance. Heqian Qiu, Lanxiao Wang, Chenghao Qi, Ruisong Dai, Hongliang Li 0001 |
MMSP | 2 |
| 2025 | GRSDet: Learning to Generate Local Reverse Samples for Few-shot Object Detection
Hefei Mei, Taijin Zhao, Shiyuan Tang, Heqian Qiu, Lanxiao Wang, Minjian Zhang 0003, Fanman Meng, Hongliang Li 0001 |
Neurocomputing | 4 |
| 2025 | Adaptively forget with crossmodal and textual distillation for class-incremental video captioning
Huiyu Xiong, Lanxiao Wang, Heqian Qiu, Taijin Zhao, Benliu Qiu, Hongliang Li 0001 |
Neurocomputing | 3 |
| 2025 | MCCE-REC: MLLM-Driven Cross-Modal Contrastive Entropy Model for Zero-Shot Referring Expression ComprehensionabstractZero-shot referring expression comprehension (zero-shot REC) is a crucial yet challenging task in the field of multi-modal understanding, which aims to locate an object described by a referring expression without training on task-specific datasets. Existing methods take advantage of a pre-trained CLIP model to align cropped proposal regions with referring expressions. However, our analysis reveals that this aligning way heavily biases toward certain salient visual regions due to CLIP focusing on global-level image-text matching. To mitigate this bias, we propose MCCE-REC, an MLLM-driven cross-modal contrastive entropy model for training-free zero-shot REC. Benefiting from the remarkable in-context comprehension ability of the multi-modal large language model (MLLM), we design a set of referring prompts for MLLM to generate diverse detailed informative, and contrastive cues related to referring objects. Based on these cues, on the one hand, we propose a multi-cues cross-modal interaction network, which associates the visual features and referring object textual features from multiple perspectives and perceives surrounding context object information in a parameter-free manner, avoiding bias towards salient features. On the other hand, we introduce a contrastive similarity entropy selection mechanism that compares the positive and negative cues to suppress biased regions with high similarity scores and emphasizes accurate regions correlating with referring descriptions. Extensive experiments demonstrate our MCCE-REC outperforms existing zero-shot methods by a significant margin on various REC datasets. Heqian Qiu, Lanxiao Wang, Taijin Zhao, Fanman Meng, Qingbo Wu 0001, Hongliang Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | Cognition Transferring and Decoupling for Text-Supervised Egocentric Semantic SegmentationabstractIn this paper, we explore a novel Text-supervised Egocentic Semantic Segmentation (TESS) task that aims to assign pixel-level categories to egocentric images weakly supervised by texts from image-level labels. In this task with prospective potential, the egocentric scenes contain dense wearer-object relations and inter-object interference. However, most recent third-view methods leverage the frozen Contrastive Language-Image Pre-training (CLIP) model, which is pre-trained on the semantic-oriented third-view data and lapses in the egocentric view due to the “relation insensitive” problem. Hence, we propose a Cognition Transferring and Decoupling Network (CTDN) that first learns the egocentric wearer-object relations via correlating the image and text. Besides, a Cognition Transferring Module (CTM) is developed to distill the cognitive knowledge from the large-scale pre-trained model to our model for recognizing egocentric objects with various semantics. Based on the transferred cognition, the Foreground-background Decoupling Module (FDM) disentangles the visual representations to explicitly discriminate the foreground and background regions to mitigate false activation areas caused by foreground-background interferential objects during egocentric relation learning. Extensive experiments on four TESS benchmarks demonstrate the effectiveness of our approach, which outperforms many recent related methods by a large margin. Code will be available athttps://github.com/ZhaofengSHI/CTDN. Zhaofeng Shi, Heqian Qiu, Lanxiao Wang, Fanman Meng, Qingbo Wu 0001, Hongliang Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Class Incremental Learning With Less Forgetting Direction and Equilibrium PointabstractCatastrophic forgetting is the core problem of class incremental learning (CIL). Existing work mainly adopts memory replay, knowledge distillation, and dynamic architecture to alleviate this problem, but seldom from the aspect of parameter regularization. However, existing parameter regularization methods struggle to achieve an appropriate balance between old and new tasks. To bring it back to CIL, we first propose constrained incremental learning with less forgetting direction (LFD) to leave more plasticity for the new task under a strong stability constraint for old tasks. Specifically, the new parameters are constrained to be close to the LFD of old tasks instead of a single group of old parameters. To validate the effectiveness of this regularization, we investigate the connectivity between the old parameters and the new parameters, and additionally find that a higher accuracy interval exists along the linear connection. Therefore, we further propose a post-processing procedure to find an equilibrium point in this interval for better balance between old and new tasks. Extensive classification experiments on CIFAR-100, ImageNet-100, and ImageNet-1K show our method can significantly improve performance compared with existing CIL methods and the object detection experiments on PASCAL-VOC show its broad generality on other tasks. Haitao Wen, Heqian Qiu, Lanxiao Wang, Haoyang Cheng, Hongliang Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Geodesic-Aligned Gradient Projection for Continual Task LearningabstractDeep networks notoriously suffer from performance deterioration on previous tasks when learning from sequential tasks, i.e., catastrophic forgetting. Recent methods of gradient projection show that the forgetting is resulted from the gradient interference on old tasks and accordingly propose to update the network in an orthogonal direction to the task space. However, these methods assume the task space is invariant and neglect the gradual change between tasks, resulting in sub-optimal gradient projection and a compromise of the continual learning capacity. To tackle this problem, we propose to embed each task subspace into a non-Euclidean manifold, which can naturally capture the change of tasks since the manifold is intrinsically non-static compared to the Euclidean space. Subsequently, we analytically derive the accumulated projection between any two subspaces on the manifold along the geodesic path by integrating an infinite number of intermediate subspaces. Building upon this derivation, we propose a novel geodesic-aligned gradient projection (GAGP) method that harnesses the accumulated projection to mitigate catastrophic forgetting. The proposed method utilizes the geometric structure information on the task manifold by capturing the gradual change between the new and the old tasks. Empirical studies on image classification demonstrate that the proposed method alleviates catastrophic forgetting and achieves on-par or better performance compared to the state-of-the-art approaches. Benliu Qiu, Heqian Qiu, Haitao Wen, Lanxiao Wang, Fanman Meng, Qingbo Wu 0001, Hongliang Li 0001 |
IEEE Trans. Image Process. | 2 |
| 2025 | Distribution-Level Memory Recall for Continual Learning: Preserving Knowledge and Avoiding ConfusionabstractContinual learning (CL) aims to enable deep neural networks (DNNs) to learn new data without forgetting previously learned knowledge. The key to achieving this goal is to avoid confusion at the feature level, i.e., to avoid confusion within old tasks and between new and old tasks. Existing prototype-based CL methods generate pseudo features for old knowledge replay by adding Gaussian noise to the centroids of old classes. However, the distribution in the feature space exhibits anisotropy during the incremental process, which prevents the pseudo features from faithfully reproducing the distribution of old knowledge in the feature space, leading to confusion at the classification boundaries within old tasks. To address this issue, we propose the distribution-level memory recall (DMR) method, which uses a Gaussian mixture model to precisely fit the feature distribution of old knowledge at the distribution level and generate pseudo features in the next stage. Furthermore, resistance to confusion at the distribution level is crucial for multimodal learning. Multimodal imbalance, which refers to uneven optimization processes among encoders of different modalities, results in significant differences in feature responses between modalities; this exacerbates confusion within old tasks in prototype-based CL methods. Therefore, we mitigate the multimodal imbalance problem by using the intermodal guidance and intramodal mining (IGIM) method to guide weaker modalities with prior information from dominant modalities and further explore useful information within modalities. To avoid confusion between new and old tasks, we propose using the confusion index to quantitatively describe a model's ability to distinguish between new and old tasks, and we use the incremental mixup feature enhancement (IMFE) method to enhance pseudo features with new sample features, alleviating classification confusion between new and old knowledge. We conduct extensive experiments on the CIFAR100, ImageNet100, TinyImageNet, ImageNet-1K and UESTC-MMEA-CL datasets and achieve state-of-the-art results. Shaoxu Cheng, Kanglei Geng, Chiyuan He, Zihuan Qiu, Linfeng Xu 0001, Heqian Qiu, Lanxiao Wang, Qingbo Wu 0001, Fanman Meng, Hongliang Li 0001 |
IEEE Trans. Multim. | 6 |
| 2024 | Prompt-Driven Referring Image Segmentation with Instance ContrastingabstractReferring image segmentation (RIS) aims to segment the target referent described by natural language. Recently, large-scale pre-trained models, e.g., CLIP and SAM, have been successfully applied in many downstream tasks, but they are not well adapted to RIS task due to inter-task differences. In this paper, we propose a new prompt-driven framework named Prompt-RIS, which bridges CLIP and SAM end-to-end and transfers their rich knowledge and powerful capabilities to RIS task through prompt learning. To adapt CLIP to pixel-level task, we first propose a Cross-Modal Prompting method, which acquires more comprehensive vision-language interaction and fine-grained text-to-pixel alignment by performing bidirectional prompting. Then, the prompt-tuned CLIP generates masks, points, and text prompts for SAM to generate more accurate mask predictions. Moreover, we further propose Instance Contrastive Learning to improve the model's discriminability to different instances and robustness to diverse languages describing the same instance. Extensive experiments demonstrate that the performance of our method outperforms the state-of-the-art methods consistently in both general and open-vocabulary settings. Chao Shang 0001, Zichen Song 0002, Heqian Qiu, Lanxiao Wang, Fanman Meng, Hongliang Li 0001 |
CVPR | 3 |
| 2024 | Class Incremental Learning with Multi-Teacher DistillationabstractDistillation strategies are currently the primary approaches for mitigating forgetting in class incremental learning (CIL). Existing methods generally inherit previous knowledge from a single teacher. However, teachers with different mechanisms are talented at different tasks, and inheriting diverse knowledge from them can enhance compatibility with new knowledge. In this paper, we propose the MTD method to find multiple diverse teachers for CIL. Specifically, we adopt weight permutation, feature perturbation, and diversity regularization techniques to ensure diverse mechanisms in teachers. To reduce time and memory consumption, each teacher is represented as a small branch in the model. We adapt existing CIL distillation strategies with MTD and extensive experiments on CIFAR-100, ImageNet-100, and ImageNet-1000 show significant performance improvement. Our code is available at https://github.com/HaitaoWen/CLearning. Haitao Wen, Lili Pan 0001, Heqian Qiu, Lanxiao Wang, Qingbo Wu 0001, Hongliang Li 0001 |
CVPR | 4 |
| 2024 | A Text Detector Based on the Specific Text PromptabstractNowadays, the prompt tuning has emerged as a novel new paradigm for adapting the original large-scale Contrastive Language-Image Pre-trained (CLIP) model into the downstream task as text detection. However, the learnable prompt adopted by the existing methods of prompt-tuning represents blurry and abstract meanings instead of fine-grained text feature. In this paper, we propose a powerful and robust text detector, called STP-TD, utilizing the specific text prompt and a learnable visual mask to fully apply the prior knowledge of CLIP model into the downstream task of text detection. STP-TD aims to make the split prompt character focusing on an ordered image token by Transformer mechanism. It is anticipated that a prompt character could stand for the fine-grained text feature of an image token, thus a distance loss is added to rectify prompt through optimization. Additionally, STP-TD firstly proposes a learnable visual mask to refine the text region in advance. Meanwhile, a synergetic framework is introduced as a bridge between visual branch and text branch. We also adopt the pixel-text matching process to align every pixel of image feature with the text feature. The experiments are conducted on the datasets ICDAR2015 and TotalText and outperform the state of the art. Xingtao Lin, Chuanyang Gong, Lanxiao Wang, Heqian Qiu, Shengyu Tong, Hongliang Li 0001 |
ICIP | 4 |
| 2024 | Video Class-Incremental Learning With Clip Based TransformerabstractVision Language Pre-training Models have shown significant potential in various domains, but there are few attempts to introduce it in the field of continual learning for video action recognition. We propose Video Class-Incremental Learner with CLIP based Transformer (VCIL-CT), which uses CLIP based vision transformer to train action recognition task by class-incremental learning pipeline. To specifically address the issue of catastrophic forgetting in transformer, we introduce Attention Distillation which distilling the attention feature from each transformer decoder. In the process of incremental learning of classes, there may be a problem of high bias towards new classes, we incorporate Class Balance Module to prevent bias on new task. Furthermore, we adopt Exemplar Augment strategy to improve exemplar quality on data replay step. We evaluate our proposed method based on the incremental action recognition benchmark presented by TCD, using UCF101, HMDB51, and UESTC-MMEA-CL datasets, and demonstrate the effectiveness of our algorithm compared to existing state-of-the-art continuous learning methods for action recognition. Shuyun Lu, Lanxiao Wang, Heqian Qiu, Xingtao Lin, Hefei Mei, Hongliang Li 0001 |
ICIP | 4 |
| 2024 | Attribute-Prompting Multi-Modal Object Reasoning Transformer for Remote Sensing Visual GroundingabstractRemote sensing visual grounding (RSVG) task aims to locate the particular object in a remote sensing image referred to a natural language expression, which requires to precisely fuse and align features from different modalities. However, existing methods usually use object-based multi-modal fusion, which is limited to capturing the detailed object characteristics in remote sensing images, resulting in object confusion with similar objects. To address this problem, we propose an attribute-prompting multi-modal object reasoning network for RSVG. Specifically, we first develop a learnable attribute prompter to adaptively explore diverse and rich attribute information according to common object characteristics in RS. With the help of attribute prompts, we design an attribute-prompting multi-modal fusion encoder to build fine-grained interactive and alignment between the visual and language features to avoid object confusion. Furthermore, we design a multi-modal progressive object reasoning decoder to gradually query more comprehensive object features for accurate object localization. Experimental results demonstrate that the proposed method achieves significant improvements. Heqian Qiu, Lanxiao Wang, Minjian Zhang 0003, Taijin Zhao, Hongliang Li 0001 |
IGARSS | 1 |
| 2024 | DP-RSCAP: Dual Prompt-Based Scene and Entity Network for Remote Sensing Image CaptioningabstractAs a challenging task towards remote sensing image analysis, the core problem of remote sensing image captioning is how to accurately transform the vision information into text information. Existing methods usually achieve it based on the simple multi-task learning strategy or visual attention mechanism, which ignores the importance of intermediate connection information for cross-modal transformation. To solve above problem, we propose a novel dual prompt-based scene and entity network (DP-RSCap) which aims to fully utilize the ability of cross-modal alignment in vision-language model build text prior information as intermediate connection to narrow the gap between different modalities and improve the quality of caption. Specifically, we first introduce an entity-concept prompt exporter to obtain explicit entity concepts in images. Then, we design a scene class prompt generator which can predict scene class and obtain fine-grained visual semantic features. Finally, we further design a dual prompt-based caption decoder to align and merge the visual semantic feature and dual prompts information as explicit intermediate connections, which can assist in generating precise caption. Extensive experiments on the challenging RSICD demonstrate the superior ability of our model. Lanxiao Wang, Heqian Qiu, Minjian Zhang 0003, Fanman Meng, Qingbo Wu 0001, Hongliang Li 0001 |
IGARSS | 2 |
| 2024 | IoU-CLIP: IoU-Aware Language-Image Model Tuning for Open Vocabulary Object DetectionabstractOpen vocabulary object detection (OVD), which detects novel categories through detectors trained on base categories, has achieved remarkable advancement attributable to large-scale vision-language models, such as CLIP. The prior OVD works mainly focused on improving the classification accuracy of proposals, ignoring the ability of localization for novel categories. In this work, we propose IoU-aware language-image model tuning (IoU-CLIP) for open vocabulary object detection. Specifically, we construct a region image dataset with different IoU and adopt IoU values as labels to fine-tune the CLIP model to learn IoU-aware and class-agnostic semantic prompts and visual embeddings. The fine-tuned IoU-CLIP can predict IoU scores for proposals, which interact with classification scores. Meanwhile, IoU-aware and class-agnostic visual embeddings are utilized for box regression to enhance the generalization of the localization capability. We evaluate our method on the COCO and LVIS OVD benchmarks, outperforming the baseline (RegionCLIP) by 5.5% AP50and 5.8% AP on novel categories, respectively, achieving state-of-the-art performance. Mingzhou He, Qingbo Wu 0001, King Ngi Ngan, Fanman Meng, Heqian Qiu, Hongliang Li 0001 |
VCIP | 6 |
| 2024 | Proposal-level Correction Guided by CLIP for Few-shot Object DetectionabstractFew-shot object detection aims at detecting previously unseen objects given only a few annotated samples. Most existing approaches treat the model obtained from the base training stage with abundant data as a container of prior knowledge that can be transferred to novel objects. Knowledge with similar properties is also contained in Contrastive Language-Image Pretraining (CLIP). In this paper, we utilize this external prior knowledge to generate proposal-level classification scores to improve the detection results. We notice that these scores can hardly reflect the quality of proposal localization, so we combine them with the ones from a conventional detector to obtain the ability to distinguish the background. Moreover, we propose a new score fusion module with regularization to alleviate the ambiguity of detection results generated by a trivial element-wise multiplication fusion method. To further improve the quality of classification scores in our proposed branch, we add learnable prompts to mitigate the inaccurate classification problem we observe. We conduct extensive experiments on the PASCAL VOC dataset and demonstrate the effectiveness of our approach. Ruihang Wang, Taijin Zhao, Hefei Mei, Heqian Qiu, Lanxiao Wang, Hongliang Li 0001 |
VCIP | 4 |
| 2024 | VLM-guided Explicit-Implicit Complementary novel class semantic learning for few-shot object detection
Taijin Zhao, Heqian Qiu, Lanxiao Wang, Hefei Mei, Fanman Meng, Qingbo Wu 0001, Hongliang Li 0001 |
Expert Syst. Appl. | 2 |
| 2024 | TridentCap: Image-Fact-Style Trident Semantic Framework for Stylized Image CaptioningabstractStylized image captioning (SIC) aims to generate captions with target style for images. The biggest challenge is that the collection and annotation of stylized data are pretty difficult and time-consuming. Most existing methods learn massive factual captions or additional stylized bookcorpus independently to assist in generating stylized caption, which ignore core relationships between existing image-fact-style trident data. In this paper, we propose a novel image-fact-style trident semantic framework TridentCap for stylized image captioning, which includes an image-fact semantic fusion encoder (SFE) and a trident stylization decoder (TSD). Unlike existing methods, we directly mine the core relationship in image-fact-style trident data and use factual semantic and image to build cross-modal semantic feature space, achieving the coherence between image and text. Specifically, SFE aims to learn the image-related prior language knowledge information from factual text and leverage fine-grained region-level semantic correlations of image and factual text to achieve cross-modal semantic information alignment and integration. TSD is designed to decouple the dual-source fused semantic feature based on the target style to achieve stylized caption generation. In addition, we design a pseudo labels filter (PLF) to obtain and expand massive image-fact-style trident data by building pseudo stylized annotations for all image-fact data in traditional caption datasets, which can further strengthen stylized caption learning. It is a generic algorithm to solve the problem of insufficient data and can be used into any existing stylized caption models. We conduct extensive experiments on SentiCap and FlickrStyle datasets, which achieve consistently improvement on almost all metrics. Our code will be released at: https://github.com/WangLanxiao/TridentCap_Code. Lanxiao Wang, Heqian Qiu, Benliu Qiu, Fanman Meng, Qingbo Wu 0001, Hongliang Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Robust Unpaired Image Dehazing via Adversarial Deformation ConstraintabstractDue to the flexible training requirement and the appealing generalization ability, unpaired image dehazing has received increasing attention in coping with real-world hazy images. However, most of the existing methods rely on the loose dehazing-hazing cycle constraint, which makes it hard to eliminate poor-quality dehazing results when using a powerful hazing network in the training process. To address this issue, this paper proposes a simple yet efficient Adversarial Deformation Constraint (ADC). More specifically, we sequentially perform two operations, i.e., dehazing and deformation, on a hazy image. In the training process, the dehazing branch is desired to be deformation-unaware, which requires that the output of these two operations remains constant regardless of their performing order. Adversarially, the deformation branch tends to maximize the difference in the outputs of these two operations when their performing orders are different. Through an additive image decomposition model, we verify that the ADC could regularize the solution space to push the dehazing error towards zero. Finally, by incorporating ADC into the common dehazing-hazing cycle constraint, we significantly improve the robustness of unpaired image dehazing. Experiments on multiple benchmark hazy image databases demonstrate the superiority of ADC over many state-of-the-art image dehazing methods. The source code of the proposed ADC-Net will be released on https://github.com/whrws/ADC-Net. Qingbo Wu 0001, King Ngi Ngan, Hongliang Li 0001, Fanman Meng, Heqian Qiu |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2024 | Continual Cross-Domain Image Compression via Entropy Prior Guided Knowledge Distillation and Scalable DecodingabstractLearning based image compression has achieved impressive rate-distortion performance in recent years. However, due to the disposable learning strategy and rigid network architecture, existing methods perform poorly for compressing the images of different domains when they emerge with the expanding real-world applications, such as, natural, oil painting, medical images and so on. To cope with this open-world challenge, this paper proposes a continual cross-domain image compression method based on entropy prior guided knowledge distillation and scalable decoding network, which perform well in balancing the plasticity, stability and compatibility. Firstly, we generate pseudo-samples of old domains by reusing their entropy priors. These pseudo-samples serve as guides for knowledge distillation in the old domains, ensuring that the bit rate and reconstruction of the new model align with those of the old model. This approach assists the updated model in retaining its capability to compress and reconstruct old images. Secondly, we develop a scalable decoding network via dynamic pruning and masked recovery, which could effectively infer an old entropy decoder from the latestly updated model. It ensures that the updated model could decode image features from binary strings encoded by old entropy encoders. Experiments on five image datasets with different domains demonstrate the effectiveness of the proposed method and its superiority over representative continual learning methods. Code of the proposed method is available athttps://github.com/wuchenhaoo/Continual_Cross-domain_Image_Compression/. Qingbo Wu 0001, Rui Ma 0030, King Ngi Ngan, Hongliang Li 0001, Fanman Meng, Heqian Qiu |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2024 | Oriented-DINO: Angle Decoupling Prediction and Consistency Optimizing for Oriented Detection TransformerabstractConsidering the arbitrary orientation of remote sensing objects, accurate angle prediction plays a crucial role in achieving precise oriented object detection (OOD) of aerial scenes. Existing transformer-based methods typically adopt an iterative refinement mechanism to update angle prediction and perform bipartite graph matching based on the combined matching costs. However, these methods may suffer from angle error accumulation across decoder layers and inconsistency between the L1 cost and the rotated intersection-of-union (IoU) cost, thus resulting in inaccurate angle prediction. To address these problems, this article proposes a novel transformer-based OOD method named Oriented-DINO (ODINO), which comprises three important components: error-mitigating angle decoupling prediction (EADP) module, nonlinear angle-conversion consistency optimizer (NACO), and query-driven diversity (QD) loss. To mitigate the angle error, the EADP module decouples angle prediction from the iterative box refinement process and uses independent branches to directly predict the angle. To address the issue of inconsistent matching, the NACO module uses a nonlinear function for angle conversion in matching cost calculation. This approach effectively alleviates the matching cost discrepancy in angle boundary case, while preserving the consistency in other instances. To avoid highly overlapped predictions triggered by similar queries, we introduce the QD loss to encourage the generation of diverse object queries, thus avoiding redundant predictions and enhancing prediction accuracy. Extensive experimental results demonstrate that our method achieves superior performance on OOD task. Minjian Zhang 0003, Heqian Qiu, Lanxiao Wang, Haoyang Cheng, Taijin Zhao, Hongliang Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Visual and Textual Prior Guided Mask Assemble for Few-Shot Segmentation and BeyondabstractFew-shot segmentation (FSS) aims to segment the novel class with a few annotated images. Due to CLIP's advantages of aligning visual and textual information, the integration of CLIP can enhance the generalization ability of FSS model. However, even with the CLIP model, the existing CLIP-based FSS methods are still subject to the biased prediction towards base class, which is caused by the class-specific feature level interactions. To solve this issue, we propose a visual and textual Prior Guided Mask Assemble Network (PGMA-Net). It employs a class-agnostic mask assembly process to alleviate the bias, and formulates diverse tasks into a unified manner by assembling the prior through affinity. Specifically, the class-relevant textual and visual features are first transformed to class-agnostic prior in the form of probability map. Then, a Prior-Guided Mask Assemble Module (PGMAM) including multiple General Assemble Units (GAUs) is introduced. It considers diverse and plug-and-play interactions, such as visual-textual, inter- and intra-image, training-free, and high-order ones. Lastly, to ensure the class-agnostic ability, a Hierarchical Decoder with Channel-Drop Mechanism (HDCDM) is proposed to flexibly exploit the assembled masks and low-level features, without relying on any class-specific information. It achieves new state-of-the-art results in the FSS task, with mIoU of 77.6 on$\rm{PASCAL-}5^{i}$and 59.4 on$\rm{COCO-}20^{i}$in 1-shot scenario. Beyond this, we show that without extra re-training, the proposed PGMA-Net can solve bbox-level and cross-domain FSS, co-segmentation, zero-shot segmentation (ZSS) tasks, leading an any-shot segmentation framework capable of accommodating diverse weak or pixel annotations. Fanman Meng, Runtong Zhang, Heqian Qiu, Hongliang Li 0001, Qingbo Wu 0001, Linfeng Xu 0001 |
IEEE Trans. Multim. | 4 |
| 2024 | CrowdCaption++: Collective-Guided Crowd Scenes CaptioningabstractCrowd scenes analysis plays an important role in various fields, including public security, smart cities, and intelligent transportation systems. However, traditional crowd scenes captioning methods mainly focus on a single and prominent crowd collective, which limits their ability to describe the different crowd collectives in complex crowd scenes. To address this issue, we propose a collective-guided crowd scenes captioning model (CrowdCaption++) to explore a more comprehensive and detailed description. We design a crowd features encoder (CFE) including double-query features encoder and foreground crowd features encoder, which uses double-query attention module (DQ-ATT) to capture more representative visual features and extracts foreground crowd features to avoid interference from background for collectives prediction. Moreover, we build a collective-guided captioning decoder (CCD) to generate captions of different crowd collectives without requiring extra alignment between crowd collectives and captions. To achieve this, we first design a crowd collectives predictor to identify multiple potential crowd collectives and create crowd collectives guidance information. Finally, we use the crowd collectives guidance information to merge useful visual features and further generate corresponding caption. We evaluate our approach on the latest crowd scenes dataset CrowdCaption and demonstrate that our model can achieve a comprehensive understanding and describe the different crowd collectives in complex crowd scenes. Lanxiao Wang, Hongliang Li 0001, Minjian Zhang 0003, Heqian Qiu, Fanman Meng, Qingbo Wu 0001, Linfeng Xu 0001 |
IEEE Trans. Multim. | 4 |
| 2024 | Learning Offset Probability Distribution for Accurate Object DetectionabstractObject detection combines object classification and object localization problems. Current object detection methods heavily depend on regression networks to locate objects, which are optimized with various regression loss functions to predict offsets between candidate boxes and objects. However, these regression losses are difficult to assign the appropriate penalties for samples with large offset errors, resulting in suboptimal regression networks and inaccurate object offsets. In this article, we consider object location as offset bin classification problem, and propose a distance-aware offset bin classification network optimized with multiple binary cross entropy losses to learn various offset probability distribution, including single label distribution and distance-aware label distribution. On one hand, it provides gradient contributions for different samples based on the bounded probability instead of previous incalculable offset error. On the other hand, it explores the distance correlations between discrete offset bins to facilitate network learning. Specifically, we discretize the continuous offset into a number of bins, and predict the probability of each offset bin, in which the probability should be higher for the offset bin closer to the target offsets, and vice versa. Furthermore, we propose an expectation-based offset prediction and a hierarchical focusing method to improve the precision of prediction. We conduct extensive experiments to evaluate the effectiveness of our method. In addition, our method can be conveniently and flexibly inserted into existing object detection methods, which consistently achieves a large gain based on popular anchor-based and anchor-free methods on the PASCAL VOC, MS-COCO, KITTI, and CrowdHuman datasets. Code will be released at: https://github.com/QiuHeqian/DBC . Heqian Qiu, Hongliang Li 0001, Qingbo Wu 0001, Hengcan Shi, Lanxiao Wang, Fanman Meng, Linfeng Xu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2023 | CafeBoost: Causal Feature Boost to Eliminate Task-Induced Bias for Class Incremental LearningabstractContinual learning requires a model to incrementally learn a sequence of tasks and aims to predict well on all the learned tasks so far, which notoriously suffers from the catastrophic forgetting problem. In this paper, we find a new type of bias appearing in continual learning, coined as task-induced bias. We place continual learning into a causal framework, based on which we find the task-induced bias is reduced naturally by two underlying mechanisms in task and domain incremental learning. However, these mechanisms do not exist in class incremental learning (CIL), in which each task contains a unique subset of classes. To eliminate the task-induced bias in CIL, we devise a causal intervention operation so as to cut off the causal path that causes the task-induced bias, and then implement it as a causal debias module that transforms biased features into unbiased ones. In addition, we propose a training pipeline to incorporate the novel module into existing methods and jointly optimize the entire architecture. Our overall approach does not rely on data replay, and is simple and convenient to plug into existing methods. Extensive empirical study on CIFAR-100 and ImageNet shows that our approach can improve accuracy and reduce forgetting of well-established methods by a large margin. Benliu Qiu, Hongliang Li 0001, Haitao Wen, Heqian Qiu, Lanxiao Wang, Fanman Meng, Qingbo Wu 0001, Lili Pan 0001 |
CVPR | 4 |
| 2023 | Incrementer: Transformer for Class-Incremental Semantic Segmentation with Knowledge Distillation Focusing on Old ClassabstractClass-incremental semantic segmentation aims to incrementally learn new classes while maintaining the capability to segment old ones, and suffers catastrophic forgetting since the old-class labels are unavailable. Most existing methods are based on convolutional networks and prevent forgetting through knowledge distillation, which (1) need to add additional convolutional layers to predict new classes, and (2) ignore to distinguish different regions corresponding to old and new classes during knowledge distillation and roughly distill all the features, thus limiting the learning of new classes. Based on the above observations, we propose a new transformer framework for class-incremental semantic segmentation, dubbed Incrementer, which only needs to add new class tokens to the transformer decoder for new-class learning. Based on the Incrementer, we propose a new knowledge distillation scheme that focuses on the distillation in the old-class regions, which reduces the constraints of the old model on the new-class learning, thus improving the plasticity. Moreover, we propose a class deconfusion strategy to alleviate the overfitting to new classes and the confusion of similar classes. Our method is simple and effective, and extensive experiments show that our method outperforms the SOTAs by a large margin (5~15 absolute points boosts on both Pascal VOC and ADE20k). We hope that our Incrementer can serve as a new strong pipeline for class-incremental semantic segmentation. Chao Shang 0001, Hongliang Li 0001, Fanman Meng, Qingbo Wu 0001, Heqian Qiu, Lanxiao Wang |
CVPR | 5 |
| 2023 | Contrastive Continuity on Augmentation Stability Rehearsal for Continual Self-Supervised LearningabstractSelf-supervised learning has attracted a lot of attention recently, which is able to learn powerful representations without any manual annotations. However, self-supervised learning needs to develop the ability to continuously learn to cope with a variety of real-world challenges, i.e., Continual Self-Supervised Learning (CSSL). Catastrophic forgetting is a notorious problem in CSSL, where the model tends to forget the learned knowledge. In practice, simple rehearsal or regularization will bring extra negative effects while alleviating catastrophic forgetting in CSSL, e.g., overfitting on the rehearsal samples or hindering the model from encoding fresh information. In order to address catastrophic forgetting without overfitting on the rehearsal samples, we propose Augmentation Stability Rehearsal (ASR) in this paper, which selects the most representative and discriminative samples by estimating the augmentation stability for rehearsal. Meanwhile, we design a matching strategy for ASR to dynamically update the rehearsal buffer. In addition, we further propose Contrastive Continuity on Augmentation Stability Rehearsal (C2ASR) based on ASR. We show that C2ASR is an upper bound of the Information Bottleneck (IB) principle, which suggests that C2ASR essentially preserves as much information shared among seen task streams as possible to prevent catastrophic forgetting and dismisses the redundant information between previous task streams and current task stream to free up the ability to encode fresh information. Our method obtains a great achievement compared with state-of-the-art CSSL methods on a variety of CSSL benchmarks. Haoyang Cheng, Haitao Wen, Xiaoliang Zhang 0002, Heqian Qiu, Lanxiao Wang, Hongliang Li 0001 |
ICCV | 4 |
| 2023 | Optimizing Mode Connectivity for Class Incremental LearningabstractClass incremental learning (CIL) is one of the most challenging scenarios in continual learning. Existing work mainly focuses on strategies like memory replay, regularization, or dynamic architecture but ignores a crucial aspect: mode connectivity. Recent studies have shown that different minima can be connected by a low-loss valley, and ensembling over the valley shows improved performance and robustness. Motivated by this, we try to investigate the connectivity in CIL and find that the high-loss ridge exists along the linear connection between two adjacent continual minima. To dodge the ridge, we propose parameter-saving OPtimizing Connectivity (OPC) based on Fourier series and gradient projection for finding the low-loss path between minima. The optimized path provides infinite low-loss solutions. We further propose EOPC to ensemble points within a local bent cylinder to improve performance on learned tasks. Our scheme can serve as a plug-in unit, extensive experiments on CIFAR-100, ImageNet-100, and ImageNet-1K show consistent improvements when adapting EOPC to existing representative CIL methods. Our code is available at https://github.com/HaitaoWen/EOPC. Haitao Wen, Haoyang Cheng, Heqian Qiu, Lanxiao Wang, Lili Pan 0001, Hongliang Li 0001 |
ICML | 3 |
| 2023 | PTCP: Alleviate Layer Collapse in Pruning at Initialization via Parameter Threshold Compensation and Preservation
Xinpeng Hao, Shiyuan Tang, Heqian Qiu, Hefei Mei, Benliu Qiu, Chuanyang Gong, Hongliang Li 0001 |
ICONIP (11) | 3 |
| 2023 | Novel-Registrable Weights and Region-Level Contrastive Learning for Incremental Few-shot Object Detection
Shiyuan Tang, Hefei Mei, Heqian Qiu, Xinpeng Hao, Taijin Zhao, Benliu Qiu, Haoyang Cheng, Chuanyang Gong, Hongliang Li 0001 |
ICONIP (11) | 3 |
| 2023 | CFS: Character Feature Summarization Model for Real-time End-to-end Text SpottingabstractMost real-time end-to-end text spotting methods employ sequence models as their recognition heads. However, these models generate characters one by one, which is inefficient when there are many characters. To solve this problem, we propose a Character Feature Summarization (CFS) Model, which can predict fixed-length characters in parallel, regardless of length. Specifically, we propose a Character Feature Summarization Module (CFSM) consisting of a Global Feature Capture and a Historical Feature Summarizer to extract and summarize global character features, enabling getting characters by simple linear prediction. We use Multi-stage Testing, cascading multiple CFSMs to obtain multi-stage summarized global character features to obtain several predictions for better convergence. The Result Selector is used to select the most likely result. Experiments on the Total-Text dataset show that CFS achieves a 3.53% improvement on the "Full" while being 3.6 times faster than ABCNet v2’s head. Chuanyang Gong, Heifei Mei, Heqian Qiu, Xinpeng Hao, Shiyuan Tang, Hongliang Li 0001 |
VCIP | 3 |
| 2023 | Disturbed Augmentation Invariance for Unsupervised Visual Representation LearningabstractContrastive learning has gained great prominence recently, which achieves excellent performance by simple augmentation invariance. However, the simple contrastive pairs suffer from lacking of diversity due to the mechanical augmentation strategies. In this paper, we propose Disturbed Augmentation Invariance (DAI for abbreviation), which constructs disturbed contrastive pairs by generating appropriate disturbed views for each augmented view in the feature space to increase the diversity. In practice, we establish a multivariate normal distribution for each augmented view, whose mean is corresponding augmented view and covariance matrix is estimated from its nearest neighbors in the dataset. Then we sample random vectors from this distribution as the disturbed views to construct disturbed contrastive pairs. In order to avoid extra computational cost with the increase of disturbed contrastive pairs, we utilize an upper bound of the trivial disturbed augmentation invariance loss to construct the DAI loss. In addition, we propose Bottleneck version of Disturbed Augmentation Invariance (BDAI for abbreviation) inspired by the Information Bottleneck principle, which further refines the extracted information and learns a compact representation by additionally increasing the variance of the original contrastive pair. In order to make BDAI work effectively, we design a statistical strategy to control the balance between the amount of the information shared by all disturbed contrastive pairs and the compactness of the representation. Our approach gets a consistent improvement over the popular contrastive learning methods on a variety of downstream tasks, e.g. image classification, object detection and instance segmentation. Haoyang Cheng, Hongliang Li 0001, Qingbo Wu 0001, Heqian Qiu, Xiaoliang Zhang 0002, Fanman Meng, Taijin Zhao |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | CrossDet++: Growing Crossline Representation for Object DetectionabstractIn object detection, precise object representation is a key factor to successfully classify and locate objects of an image. Existing methods usually use rectangular anchor boxes or a set of points to represent objects. However, these methods either introduce background noise or miss the continuous appearance information inside the object, and thus cause incorrect detection results. In this paper, we propose a novel anchor-free object detection network, called CrossDet++, which uses a set of growing crosslines along horizontal and vertical axes as object representations. An object can be flexibly represented as crosslines in different combinations, which inspires us to select the expressive crossline to effectively reduce the interference of noise. Meanwhile, the crossline representation takes into account the continuous adjacent object information, which is useful to enhance the discriminability of object features and find the object boundaries. Based on the learned crosslines, we propose an axis-query crossline growing module to adaptively capture features of crosslines and query surrounding pixels related to the line features for subsequent growing of crosslines. Their growing offsets and scales can be supervised by a decoupled regression mechanism, which limits the regression target to a specific direction for decreasing the optimization difficulty. During the training, we design a semantic-guided label assignment to emphasize the importance of crossline targets with higher semantic richness, further improving the detection performance. The experiment results demonstrate the effectiveness of our proposed method. Code can be available at:https://github.com/QiuHeqian/CrossDet. Heqian Qiu, Hongliang Li 0001, Qingbo Wu 0001, Jianhua Cui, Zichen Song 0002, Lanxiao Wang, Minjian Zhang 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | Cross-Modal Recurrent Semantic Comprehension for Referring Image SegmentationabstractReferring image segmentation aims to segment the target object from the image according to the description of language expression. Due to the diversity of language expressions, word sequences in different orders often express different semantic information. The previous methods focus more on matching different words to different visual regions in the image separately, ignoring the global semantic understanding of language expression based on the sequence structure. To address this problem, we redesign a new recurrent network structure for referring image segmentation, called Cross-Modal Recurrent Semantic Comprehension Network (CRSCNet), to obtain a more comprehensive global semantic understanding through iterative cross-modal semantic reasoning. Specifically, in each iteration, we first propose a Dynamic SepConv to extract relevant visual features guided by language and further propose Language Attentional Feature Modulation to improve the feature discriminability, then propose a Cross-Modal Semantic Reasoning module to perform global semantic reasoning by capturing both linguistic and visual information, and finally updates and corrects the visual features of the predicted object based on semantic information. Moreover, we further propose a Cross-Modal ASPP to capture richer visual information referred to in the global semantics of the language expression from larger receptive fields. Extensive experiments demonstrate that our proposed network significantly outperforms previous state-of-the-art methods on multiple datasets. Chao Shang 0001, Hongliang Li 0001, Heqian Qiu, Qingbo Wu 0001, Fanman Meng, Taijin Zhao, King Ngi Ngan |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | DRDet: Dual-Angle Rotated Line Representation for Oriented Object DetectionabstractIn aerial scenes, oriented object detection is sensitive to the orientation of objects, which makes the formulation of orientation-aware object representation become a critical problem. Existing methods mostly adopt rectangle anchor or discrete points as object representation, which may lead to the feature aliasing between overlapping objects and ignore the orientation information of objects. To solve these issues, we propose a novel anchor-free oriented object detection network named DRDet, which adopts Dual-angle Rotated Lines (DRL) as object representation. Different from other object representations, DRL can adaptively rotate and extend to the boundary of the object according to its orientation and shape, which explicitly introduces the orientation information into the formulation of object representation. And it can adaptively cope with the geometric deformation of objects. Based on the dual-angle rotated lines, we design an Orientation-guided Feature Encoder (OFE) to encode discriminant object feature along each rotated line, respectively. Instead of encoding rectangle feature, the OFE module adopts line features for orientation-guided feature encoding, which can alleviate the feature aliasing between neighboring objects or background. To further enhance the flexibility of dual-angle rotated lines, we design a Dual-angle Decoder (DD) that predicts two angle offsets according to the orientation-guided feature and converts the angle offsets and regression offsets into dual-angle rotated line representation, which can help to guide the adaptive rotation of each rotated line, respectively. Our proposed method achieves consistent improvement on both DOTA and HRSC2016 datasets. Extensive experimental results verify the effectiveness of our method in oriented object detection. Minjian Zhang 0003, Heqian Qiu, Hefei Mei, Lanxiao Wang, Fanman Meng, Linfeng Xu 0001, Hongliang Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Unsupervised Visual Representation Learning via Multi-Dimensional Relationship AlignmentabstractRecently, contrastive learning based on augmentation invariance and instance discrimination has made great achievements, owing to its excellent ability to learn beneficial representations without any manual annotations. However, the natural similarity among instances conflicts with instance discrimination which treats each instance as a unique individual. In order to explore the natural relationship among instances and integrate it into contrastive learning, we propose a novel approach in this paper, Relationship Alignment (RA for abbreviation), which forces different augmented views of current batch instances to main a consistent relationship with other instances. In order to perform RA effectively in existing contrastive learning framework, we design an alternating optimization algorithm where the relationship exploration step and alignment step are optimized respectively. In addition, we add an equilibrium constraint for RA to avoid the degenerate solution, and introduce the expansion handler to make it approximately satisfied in practice. In order to better capture the complex relationship among instances, we additionally propose Multi-Dimensional Relationship Alignment (MDRA for abbreviation), which aims to explore the relationship from multiple dimensions. In practice, we decompose the final high-dimensional feature space into a cartesian product of several low-dimensional subspaces and perform RA in each subspace respectively. We validate the effectiveness of our approach on multiple self-supervised learning benchmarks and get consistent improvements compared with current popular contrastive learning methods. On the most commonly used ImageNet linear evaluation protocol, our RA obtains significant improvements over other methods, our MDRA gets further improvements based on RA to achieve the best performance. The source code of our approach will be released soon. Haoyang Cheng, Hongliang Li 0001, Heqian Qiu, Qingbo Wu 0001, Xiaoliang Zhang 0002, Fanman Meng, King Ngi Ngan |
IEEE Trans. Image Process. | 3 |
| 2023 | What Happens in Crowd Scenes: A New Dataset About Crowd Scenes for Image CaptioningabstractMaking machines endowed with eyes and brains to effectively understand and analyze crowd scenes is of paramount importance for building a smart city to serve people. This is of far-reaching significance for the guidance of dense crowds and accident prevention, such as crowding and stampedes. As a typical multimodal scene understanding task, image captioning has always attracted widespread attention. However, crowd scene understanding captioning is rarely studied due to the unobtainability of related datasets. Therefore, it is difficult to know what happens in crowd scenes. In order to fill this research gap, we propose a crowd scenes caption dataset named CrowdCaption which has the advantages of crowd-topic scenes, comprehensive and complex caption descriptions, typical relationships and detailed grounding annotations. The complexity and diversity of the descriptions and the specificity of the crowd scenes make this dataset extremely challenging to most current methods. Thus, we propose a Multi-hierarchical Attribute Guided Crowd Caption Network (MAGC) based on crowd objects, actions, and status (such as position, dress, posture, etc.) aiming to generate crowd-specific detailed descriptions. We conduct extensive experiments on our CrowdCaption dataset, and our proposed method reaches the state-of-the-art (SoTA) performance. We hope the CrowdCaption dataset can assist future studies related to crowd scenes in the multimodal domain. Lanxiao Wang, Hongliang Li 0001, Wenzhe Hu, Xiaoliang Zhang 0002, Heqian Qiu, Fanman Meng, Qingbo Wu 0001 |
IEEE Trans. Multim. | 5 |
| 2023 | Bias-Correction Feature Learner for Semi-Supervised Instance SegmentationabstractInstance segmentation is heavily reliant on large-scale annotated datasets to yield an ideal accuracy. However, annotated data are difficult to collect. To expand the annotated data, a straightforward idea is to introduce semi-supervised learning, which uses a trained model to obtain initial proposals on unlabeled images and then use initial proposals to generate pseudo labels. However, existing methods inevitably introduce the bias for the model learning, i.e., the foreground in initial low-confident proposals (low-confident foreground) is arbitrarily assigned as background. This bias makes the foreground and background closer in the feature space, which degenerates the model accuracy. To address this issue, this paper discards incorrect supervision and designs a bias-correction feature learner. Specifically, on the one hand, low-confident foreground does not participate in supervised learning. On the other hand, we extract possible foreground regions from all initial proposals to construct high-quality positive pairs which depict objects of the same category in contrastive learning. Then, positive pairs are pulled closer in the feature space. This helps models extract closely clustered foreground features. Experimental results demonstrate the effectiveness of our method on the public datasets (i.e., COCO, Cityscapes and Pascal VOC). Longrong Yang, Hongliang Li 0001, Qingbo Wu 0001, Fanman Meng, Heqian Qiu, Linfeng Xu 0001 |
IEEE Trans. Multim. | 5 |
| 2022 | RefCrowd: Grounding the Target in Crowd with Referring ExpressionsabstractCrowd understanding has aroused the widespread interest in vision domain due to its important practical significance. Unfortunately, there is no effort to explore crowd understanding in multi-modal domain that bridges natural language and computer vision. Referring expression comprehension (REF) is such a representative multi-modal task. Current REF studies focus more on grounding the target object from multiple distinctive categories in general scenarios. It is difficult to applied to complex real-world crowd understanding. To fill this gap, we propose a new challenging dataset, called RefCrowd, which towards looking for the target person in crowd with referring expressions. It not only requires to sufficiently mine natural language information, but also requires to carefully focus on subtle differences between the target and a crowd of persons with similar appearance, so as to realize fine-grained mapping from language to vision. Furthermore, we propose a Fine-grained Multi-modal Attribute Contrastive Network (FMAC) to deal with REF in crowd understanding. It first decomposes the intricate visual and language features into attribute-aware multi-modal features, and then captures discriminative but robustness fine-grained attribute features to effectively distinguish these subtle differences between similar persons. The proposed method outperforms existing state-of-the-art (SoTA) methods on our RefCrowd dataset and existing REF datasets. In addition, we implement an end-to-end REF toolbox for the deeper research in multi-modal domain. Our dataset and code can be available at: https://qiuheqian.github.io/datasets/refcrowd/. Heqian Qiu, Hongliang Li 0001, Taijin Zhao, Lanxiao Wang, Qingbo Wu 0001, Fanman Meng |
ACM Multimedia | 1 |
| 2022 | Cross-Domain Object Detection with Missing Classes in Target DomainabstractMany existing methods focus on detecting either objects from different domains or those of rare classes, but it's difficult for them to tackle the two issues together. However, in the real world, due to the difficulty of collecting samples of special classes, deep learning practitioners have to use simulated images to substitute for them. To deal with this scenario, in this paper, we research a new task: cross-domain object detection with missing classes in target domain, where there are only partial classes have images and annotations in the target domain. We devise a simple but effective play-and-plug method to address this new task, named the three-stage learning approach with domain and class information preservation. In addition, extensive experiments demonstrate our method is effective and can boost the performance when added to existing unsupervised domain adaptation object detectors. Benliu Qiu, Heqian Qiu, Haitao Wen, Zichen Song 0002, Linfeng Xu 0001 |
MMSP | 2 |
| 2022 | DE-CrossDet: Divisible and Extensible Crossline Representation for Object DetectionabstractObject detection aims to localize and classify objects. Suitable object representation plays an important role in accurate detection. Because a complete crossline inevitably passes through the noise of backgrounds or other objects, object features directly extracted by the whole crossline are often confused. In this paper, we present a new feature extraction method, DE-Crossline, which can enhance the original crossline representation to capture more accurate object information. Specifically, we divide the crossline into several segments, each of which extracts the maximum activation key point respectively to reduce the impact of noise mentioned above. Furthermore, considering various shapes and sizes of objects, we design a Deformable Width Extension Module to learn a suitable width of each crossline, so as to capture richer object information. Extensive experiments prove the effectiveness of our proposed method. The total performance of our proposed detector can reach 49.0% AP, using ResNet-101 as backbone on the MS-COCO dataset. Hefei Mei, Hongliang Li 0001, Heqian Qiu, Jianhua Cui, Longrong Yang |
VCIP | 3 |
| 2022 | Mining Regional Relation from Pixel-wise Annotation for Scene ParsingabstractScene parsing is an important and challenging task in computer vision, which assigns semantic labels to each pixel in the entire scene. Existing scene parsing methods only utilize pixel-wise annotation as the supervision of neural network, thus, some similar categories are easy to be misclassified in the complex scenes without the utilization of regional relation. To tackle these above challenging problems, a Regional Relation Network (RRNet) is proposed in this paper, which aims to boost the scene parsing performance by mining regional relation from pixel-wise annotation. Specifically, the pixel-wise annotation is divided into a lot of fixed regions, so that intra- and inter-regional relation are able to be extracted as the supervision of network. We firstly design an intra-regional relation module to predict category distribution in each fixed region, which is helpful for reducing the misclassification phenomenon in regions. Secondly, an inter-regional relation module is proposed to learn the relationships among each region in scene images. With the guideline of relation information extracted from the ground truth, the network is able to learn more discriminative relation representations. To validate our proposed model, we conduct experiments on three typical datasets, including NYU-depth-v2, PASCAL-Context and ADE20k. The achieved competitive results on all three datasets demonstrate the effectiveness of our method. Zichen Song 0002, Hongliang Li 0001, Heqian Qiu, Xiaoliang Zhang 0002 |
VCIP | 3 |
| 2022 | Instance-level Context Attention Network for instance segmentation
Chao Shang 0001, Hongliang Li 0001, Fanman Meng, Heqian Qiu, Qingbo Wu 0001, Linfeng Xu 0001, King Ngi Ngan |
Neurocomputing | 4 |
| 2022 | Real-time panoptic segmentation with relationship between adjacent pixels and boundary prediction
Xiaoliang Zhang 0002, Hongliang Li 0001, Lanxiao Wang, Haoyang Cheng, Heqian Qiu, Wenzhe Hu, Fanman Meng, Qingbo Wu 0001 |
Neurocomputing | 5 |
| 2022 | POS-Trends Dynamic-Aware Model for Video CaptionabstractVideo caption aims to generate descriptive sentences about the video, and the most critical problem is how to achieve accurate word prediction with standardized and coherent syntax structure, which requires the model to thoroughly understand video content and precisely map them into corresponding sentence components. Many existing methods usually fuse different video features into a single visual feature for generating sentences. However, they ignore the word dataset prior information in the annotations (such as Part-Of-Speech) and they also ignore the association between sentence components and types of visual features. To solve these problems, we propose a POS-trends dynamic-aware model (PDA) to fully exploit the word dataset prior information in the captions to predict POS tag, so as to assist generating captions. We propose a POS feature extraction (PFE) module to use different filters to extract different POS-trends features, predict POS tags and fuse visual features. Furthermore, we propose a visual-dynamic-aware (VDA) module to dynamically adjust the mapping way of words and supplement the visual information into the local features. The fusion features provide directional visual information to generate correct words, and the predicted POS tags to guide the decoding process to generate a more standardized and coherent syntax structure. A large number of experiments based on MSVD, MSR-VTT and VATEX demonstrated that our method outperforms the state-of-the-art methods in BLEU-4, ROUGE-L, METEOR, CIDEr. Code can be available at:https://github.com/WangLanxiao/PDA-for-video-caption. Lanxiao Wang, Hongliang Li 0001, Heqian Qiu, Qingbo Wu 0001, Fanman Meng, King Ngi Ngan |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Bal-R$^2$CNN: High Quality Recurrent Object Detection With Balance OptimizationabstractIt is a common practice to refine object detection results using recurrent detection paradigm. We evaluate the recurrent detection on Faster R-CNN, but the improvement is far away from expected. We consider that the performance bottleneck is fromimbalance optimizationcaused by the biased distribution of training data. Low-IoU-skewed RPN proposals could suppress the contribution of High-IoU examples at the training stage. Besides, data imbalance and statistical discrepancy on regression targets between low-IoU and high-IoU examples are not considered in the regression task; this design could impede localization quality. In this work, we propose Bal-R$^2$CNN for high-quality recurrent object detection. There are two new components in Bal-R$^2$CNN.Self-iteration box samplingcollects object boxes from recurrent steps and increases the number of high-IoU training examples.IoU-sensitive bounding-box regressionsends proposal boxes with different IoUs to specified regression branches for more accurate bounding-box prediction. Both two new components could inducebalanced optimizationand be helpful. With the resulting Bal-R$^2$CNN detector, evaluation on PASCAL VOC and MSCOCO reveal that our method has a significant improvement on the existing solution and could reach a better performance than several state-of-the-art methods. Hongliang Li 0001, Qingbo Wu 0001, Fanman Meng, Heqian Qiu |
IEEE Trans. Multim. | 5 |
| 2021 | CrossDet: Crossline Representation for Object DetectionabstractObject detection aims to accurately locate and classify objects in an image, which requires precise object representations. Existing methods usually use rectangular anchor boxes or a set of points to represent objects. However, these methods either introduce background noise or miss the continuous appearance information inside the object, and thus cause incorrect detection results. In this paper, we propose a novel anchor-free object detection network, called Cross-Det, which uses a set of growing cross lines along horizontal and vertical axes as object representations. An object can be flexibly represented as cross lines in different combinations. It not only can effectively reduce the interference of noise, but also take into account the continuous object information, which is useful to enhance the discriminability of object features and find the object boundaries. Based on the learned cross lines, we propose a crossline extraction module to adaptively capture features of cross lines. Furthermore, we design a decoupled regression mechanism to regress the localization along the horizontal and vertical directions respectively, which helps to decrease the optimization difficulty because the optimization space is limited to a specific direction. Our method achieves consistently improvement on the PASCAL VOC and MS-COCO datasets. The experiment results demonstrate the effectiveness of our proposed method. Code can be available at: https://github.com/QiuHeqian/CrossDet. Heqian Qiu, Hongliang Li 0001, Qingbo Wu 0001, Jianhua Cui, Zichen Song 0002, Lanxiao Wang, Minjian Zhang 0003 |
ICCV | 1 |
| 2020 | Offset Bin Classification Network for Accurate Object DetectionabstractObject detection combines object classification and object localization problems. Most existing object detection methods usually locate objects by leveraging regression networks trained with Smooth L1loss function to predict offsets between candidate boxes and objects. However, this loss function applies the same penalties on different samples with large errors, which results in suboptimal regression networks and inaccurate offsets. In this paper, we propose an offset bin classification network optimized with cross entropy loss to predict more accurate offsets. It not only provides different penalties for different samples but also avoids the gradient explosion problem caused by the samples with large errors. Specifically, we discretize the continuous offset into a number of bins, and predict the probability of each offset bin. Furthermore, we propose an expectation-based offset prediction and a hierarchical focusing method to improve the prediction precision. Extensive experiments on the PASCAL VOC and MS-COCO datasets demonstrate the effectiveness of our proposed method. Our method outperforms the baseline methods by a large margin. Heqian Qiu, Hongliang Li 0001, Qingbo Wu 0001, Hengcan Shi |
CVPR | 1 |
| 2020 | Language-Aware Fine-Grained Object Representation for Referring Expression ComprehensionabstractReferring expression comprehension expects to accurately locate an object described by a language expression, which requires precise language-aware visual object representations. However, existing methods usually use rectangular object representations, such as object proposal regions and grid regions. They ignore some fine-grained object information like shapes and poses, which are often described in language expressions and important to localize objects. Additionally, rectangular object regions usually contain background contents and irrelevant foreground features, which also decrease the localization performance. To address these problems, we propose a language-aware deformable convolution model (LDC) to learn language-aware fine-grained object representations. Rather than extracting rectangular object representations, LDC adaptively samples a set of key points based on the image and language to represent objects. This type of object representations can capture more fine-grained object information (e.g., shapes and poses) and suppress noises in accordance with language and thus, boosts the object localization performance. Based on the language-aware fine-grained object representation, we next design a bidirectional interaction model (BIM) that leverages a modified co-attention mechanism to build cross-modal bidirectional interactions to further improve the language and object representations. Furthermore, we propose a hierarchical fine-grained representation network (HFRN) to learn language-aware fine-grained object representations and cross-modal bidirectional interactions at local word level and global sentence level, respectively. Our proposed method outperforms the state-of-the-art methods on the RefCOCO, RefCOCO+ and RefCOCOg datasets. Heqian Qiu, Hongliang Li 0001, Qingbo Wu 0001, Fanman Meng, Hengcan Shi, Taijin Zhao, King Ngi Ngan |
ACM Multimedia | 1 |
| 2020 | Multi-stage Tag Guidance Network in Video CaptionabstractRecently, video caption plays an important role in computer vision tasks. We participate in Pre-training for Video Captioning Challenge which aims to produce at least one sentence for each challenge video based on the pretraining models. In this work, we propose a tag guidance module to learn a representation which can better build the interaction in cross-modal between visual content and textual sentences. First, we utilize three types of features extraction networks to fully capture the information of 2D, 3D and object information. Second, to prevent overfitting and time issues, the entire process of training is divided into two stages. The first stage trains all data, and the second stage introduces a random dropout. Furthermore, we train a CNN-based network to pick out the best candidate results. In summary, we were ranked third place in Pre-training for Video Captioning Challenge which proved the effectiveness of our model. Lanxiao Wang, Chao Shang 0001, Heqian Qiu, Taijin Zhao, Benliu Qiu, Hongliang Li 0001 |
ACM Multimedia | 3 |
| 2020 | A multi-scale language embedding network for proposal-free referring expression comprehensionabstractReferring expression comprehension (REC) is a task that aims to find the location of an object specified by a language expression. Current solutions for REC can be classified into proposal-based methods and proposal-free methods. Proposal-free methods are popular recently because of its flexibility and lightness. Nevertheless, existing proposal-free works give little consideration to visual context. As REC is a context sensitive task, it is hard for current proposal-free methods to comprehend expressions that describe objects by the relative position with surrounding things. In this paper, we propose a multi-scale language embedding network for REC. Our method adopts the proposal-free structure, which directly feeds fused visual-language features into a detection head to predict the bounding box of the target. In the fusion process, we propose a grid fusion module and a grid-context fusion module to compute the similarity between language features and visual features in different size regions. Meanwhile, we extra add fully interacted vision-language information and position information to strength the feature fusion. This novel fusion strategy can help to utilize context flexibly therefore the network can deal with varied expressions, especially expressions that describe objects by things around. Our proposed method outperforms the state-of-the-art methods on Refcoco, Refcoco+ and Refcocog datasets. Taijin Zhao, Hongliang Li 0001, Heqian Qiu, Qingbo Wu 0001, King Ngi Ngan |
MMAsia | 3 |
| 2020 | Hierarchical Context Features Embedding for Object DetectionabstractPixel-level segmentation has been widely used to improve object detection. Most of the existing methods refine detection features by adding the constraint of the segmentation branch or by simply embedding high-level segmentation features into detection features within the local receptive field. However, noisy segmentation features are unavoidable in real-word applications and can easily cause false positives. To address this problem, we propose a novel hierarchical context embedding module to effectively embed segmentation features into detection features. The idea of this module is to capture hierarchical context information that includes local objects or parts and nonlocal context features by learning multiple attention maps, and subsequently utilize interdependencies between features to recalibrate noisy segmentation features. Furthermore, we use this module in the proposed gated encoder-decoder network that adaptively aggregates feature maps of different resolutions based on the gate mechanism so that we can embed multiscale segmentation feature maps into detection features for more accurate detection of objects of all sizes. Experimental results demonstrate the effectiveness of the proposed method on the Pascal VOC 2012Seg dataset, the Pascal VOC dataset and the MS COCO dataset. Heqian Qiu, Hongliang Li 0001, Qingbo Wu 0001, Fanman Meng, Linfeng Xu 0001, King Ngi Ngan, Hengcan Shi |
IEEE Trans. Multim. | 1 |