VLDB 2026 Research / reviewers in the wild / expert
Lanxiao Wang
dblp:276/3237
· DBLP profile ↗
39ranked-venue papers
7as first author
38since 2021 · last 2026
0000-0002-3745-0262ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 29 · 5 first-author · 28 since 2021Artificial intelligence and machine learning · 14 · 1 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Parameter Merging with Gradient-Guided Supermasks in Online Continual LearningabstractOnline continual learning (OCL) aims at learning a non-stationary data stream in a way of reading each data sample only once, and hence suffers from the trade-off of catastrophic forgetting and insufficient learning. In this work, we firstly analytically establish relationship between loss functions and model parameters from the Bayesian perspective. Based on our analysis, we subsequently propose a parameter merging method with gradient-guided supermasks. Our method leverages 1-order and 2-order gradient information to construct supermasks that determine the merging weights between the old and new models. Our method performs direct arithmetic operations on parameters to update models, beyond traditional gradient descent. We further discover that a widely-used premise that 1-order gradients can be negligible is invalid in OCL, due to slow convergence incurred by insufficient learning. Additionally, we utilize a dual-model dual-view distillation strategy that can align output distributions of the new and merged models for each sample, further enhancing model performance. Extensive experiments are conducted on four benchmarks in OCL settings, including CIFAR-10, CIFAR-100, Tiny-ImageNet, and ImageNet-100. Experimental results demonstrate that our method is effective, and achieves a substantial boost over previous methods. Benliu Qiu, Heqian Qiu, Lanxiao Wang, Taijin Zhao, Lili Pan 0001, Hongliang Li 0001 |
AAAI | 3 |
| 2026 | Ego-PMOVE: Prompt-aware Mixture of View Experts Network for Egocentric Gaze PredictionabstractEgocentric gaze prediction serves as a critical indicator for decoding human visual attention and cognitive processes, but its inherently limited field of view creates prediction challenges. Although exo-view data provides supplementary contextual information, it exhibits significant spatial and semantic gaps. Existing methods focus solely on isolated feature encoding in single-view paradigms, neglecting cross-view gaze correlations. To make up for this gap, we make the first exploration of cross-view gaze relationship for egocentric gaze prediction, and propose Ego-PMOVE, a novel Prompt-aware Mixture of View Experts network. Unlike prior cross-view studies that forcibly align cross-view features thereby introducing inference noise, we leverage the popular Mixture-of-Experts (MoE) and a set of flexible prompts to disentangle features from different views into three parallel experts: a view-shared expert directly modeling common semantic relationships, a view-discrepancy expert adaptively adjusting the spatial position, scale and shifts based on different view-specific features, and an egocentric expert extracting independent features to compensate for the case of missing exocentric data. To balance these experts, we further design a soft router to dynamically weight them for mining useful information while suppressing noise. A view-query gaze decoder then generates view-specific gaze attention maps, jointly optimized by gaze-heamap and cross-view contrastive loss that regularize both shared and divergent features for accurate gaze prediction. Extensive experiments across the multi-view EgoMe dataset and single-view Ego4D and EGTEA Gaze++ datasets demonstrate the effectiveness and generalizability of our approach. Heqian Qiu, Lanxiao Wang, Taijin Zhao, Zhaofeng Shi, Linfeng Xu 0001, Hongliang Li 0001 |
AAAI | 2 |
| 2026 | Bridging the Gap between Vision and Text for unsupervised text-only captioning
Lanxiao Wang, Heqian Qiu, Haitao Wen, Fanman Meng, Qingbo Wu 0001, Hongliang Li 0001 |
Pattern Recognit. | 1 |
| 2025 | Unsupervised Ego- and Exo-centric Dense Procedural Activity Captioning via Gaze Consensus AdaptationabstractEven from an early age, humans naturally adapt between exocentric (Exo) and egocentric (Ego) perspectives to understand daily procedural activities. Inspired by this cognitive ability, we propose a novel Unsupervised Ego-Exo Dense Procedural Activity Captioning (UE^2 DPAC) task, which aims to transfer knowledge from the labeled source view to predict the time segments and descriptions of action sequences for the target view without annotations. Despite previous works endeavoring to address the fully-supervised single-view or cross-view dense video captioning, they lapse in the proposed task due to the significant inter-view gap caused by temporal misalignment and irrelevant object interference. Hence, we propose a Gaze Consensus-guided Ego-Exo Adaptation Network (GCEAN) that injects the gaze information into the learned representations for the fine-grained Ego-Exo alignment. Specifically, we propose a Score-based Adversarial Learning Module (SALM) that incorporates a discriminative scoring network and compares the scores of distinct views to learn unified view-invariant representations from a global level. Then, the Gaze Consensus Construction Module (GCCM) utilizes the gaze to progressively calibrate the learned representations to highlight the regions of interest and extract the corresponding temporal contexts. Moreover, we adopt hierarchical gaze-guided consistency losses to construct gaze consensus for the explicit temporal and spatial adaptation between the source and target views. To support our research, we propose a new EgoMe-UE^2 DPAC benchmark, and extensive experiments demonstrate the effectiveness of our method, which outperforms many related methods by a large margin. Code is available at https://github.com/ZhaofengSHI/GCEAN. Zhaofeng Shi, Heqian Qiu, Lanxiao Wang, Qingbo Wu 0001, Fanman Meng, Hongliang Li 0001 |
ACM Multimedia | 3 |
| 2025 | D3Net: Dual-Path Decoupling-Distillation for Adaptive Fusion in Continual Egocentric LearningabstractEgocentric continual action recognition faces severe challenges such as sudden viewpoint changes, occlusions, and complex backgrounds. In such scenarios, relying solely on visual modalities is susceptible to interference and lacks sufficient recognition robustness. To overcome the limitations of unimodal approaches, multimodal fusion methods are widely adopted, significantly enhancing recognition performance. However, existing multimodal schemes generally suffer from insufficient exploration of cross-modal complementarity and the vulnerability of modal independence. To address this, this paper proposes a Dual-path Decoupling-Distillation NetWork (D3Net), aiming to achieve more effective dynamic fusion of modal information and knowledge transfer.D3Net first explicitly separates the shared and private features of modalities through a dual-path decoupling module, combined with a dynamic gating mechanism to adaptively adjust the modal fusion weights. Secondly, it designs a complementary distillation module, leveraging cross-modal contrastive learning to effectively mitigate the issues of poor unimodal robustness and vulnerability to interference. Finally, through a cross-task distillation mechanism, it efficiently extracts knowledge from old tasks, alleviating the catastrophic forgetting problem during learning. Experimental results demonstrate that D3Net achieves an average accuracy of 83.97% under the 8×4 task configuration on the UESTC MMEA CL dataset, surpassing baseline method by 5.17%. Chenghao Qi, Heqian Qiu, Zhaofeng Shi, Lanxiao Wang, Hongliang Li 0001 |
MMSP | 4 |
| 2025 | Efficient Polyp Detection via Wavelet-Driven Boundary Enhancement and Temporal ConsistencyabstractAccurate early detection of polyps plays a critical role in preventing, diagnosing, and treating colorectal cancer. Although significant progress has been made, the accurate and efficient detection of polyps remains a challenging task. Existing image-based methods, while computationally efficient, typically rely on single-frame inputs and struggle to handle polyps with ambiguous boundaries or varying sizes. On the other hand, video-based approaches leverage temporal information to improve detection robustness, but often incur high computational costs and compromise real-time performance due to the need to process multiple frames simultaneously. Moreover, both types of methods are susceptible to dynamic artifacts caused by endoscopic camera movement, which can lead to polyp-like false positives. To address these issues, we propose BEC-Net, a novel Boundary-Enhanced network with adjacent-Frame Contrastive Learning for accurate and efficient polyp detection. Specifically, we design a Wavelet-Based Boundary-Aware feature Fusion (WBAF) module to enhance the representation of polyp boundaries and improve generalization across diverse appearances. To accommodate scale variation, we introduce a Context-Aware Gated Aggregation (CAGA) module that adaptively integrates multi-scale contextual information. Furthermore, we propose an Adjacent-Frame Contrastive Learning (AFCL) strategy that utilizes temporal consistency between adjacent frames to suppress polyp-like artifacts without increasing inference cost. Extensive experiments on large-scale colonoscopy benchmarks demonstrate that our method outperforms state-of-the-art approaches in both accuracy and real-time performance. Heqian Qiu, Lanxiao Wang, Chenghao Qi, Ruisong Dai, Hongliang Li 0001 |
MMSP | 3 |
| 2025 | GRSDet: Learning to Generate Local Reverse Samples for Few-shot Object Detection
Hefei Mei, Taijin Zhao, Shiyuan Tang, Heqian Qiu, Lanxiao Wang, Minjian Zhang 0003, Fanman Meng, Hongliang Li 0001 |
Neurocomputing | 5 |
| 2025 | Adaptively forget with crossmodal and textual distillation for class-incremental video captioning
Huiyu Xiong, Lanxiao Wang, Heqian Qiu, Taijin Zhao, Benliu Qiu, Hongliang Li 0001 |
Neurocomputing | 2 |
| 2025 | MCCE-REC: MLLM-Driven Cross-Modal Contrastive Entropy Model for Zero-Shot Referring Expression ComprehensionabstractZero-shot referring expression comprehension (zero-shot REC) is a crucial yet challenging task in the field of multi-modal understanding, which aims to locate an object described by a referring expression without training on task-specific datasets. Existing methods take advantage of a pre-trained CLIP model to align cropped proposal regions with referring expressions. However, our analysis reveals that this aligning way heavily biases toward certain salient visual regions due to CLIP focusing on global-level image-text matching. To mitigate this bias, we propose MCCE-REC, an MLLM-driven cross-modal contrastive entropy model for training-free zero-shot REC. Benefiting from the remarkable in-context comprehension ability of the multi-modal large language model (MLLM), we design a set of referring prompts for MLLM to generate diverse detailed informative, and contrastive cues related to referring objects. Based on these cues, on the one hand, we propose a multi-cues cross-modal interaction network, which associates the visual features and referring object textual features from multiple perspectives and perceives surrounding context object information in a parameter-free manner, avoiding bias towards salient features. On the other hand, we introduce a contrastive similarity entropy selection mechanism that compares the positive and negative cues to suppress biased regions with high similarity scores and emphasizes accurate regions correlating with referring descriptions. Extensive experiments demonstrate our MCCE-REC outperforms existing zero-shot methods by a significant margin on various REC datasets. Heqian Qiu, Lanxiao Wang, Taijin Zhao, Fanman Meng, Qingbo Wu 0001, Hongliang Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Cognition Transferring and Decoupling for Text-Supervised Egocentric Semantic SegmentationabstractIn this paper, we explore a novel Text-supervised Egocentic Semantic Segmentation (TESS) task that aims to assign pixel-level categories to egocentric images weakly supervised by texts from image-level labels. In this task with prospective potential, the egocentric scenes contain dense wearer-object relations and inter-object interference. However, most recent third-view methods leverage the frozen Contrastive Language-Image Pre-training (CLIP) model, which is pre-trained on the semantic-oriented third-view data and lapses in the egocentric view due to the “relation insensitive” problem. Hence, we propose a Cognition Transferring and Decoupling Network (CTDN) that first learns the egocentric wearer-object relations via correlating the image and text. Besides, a Cognition Transferring Module (CTM) is developed to distill the cognitive knowledge from the large-scale pre-trained model to our model for recognizing egocentric objects with various semantics. Based on the transferred cognition, the Foreground-background Decoupling Module (FDM) disentangles the visual representations to explicitly discriminate the foreground and background regions to mitigate false activation areas caused by foreground-background interferential objects during egocentric relation learning. Extensive experiments on four TESS benchmarks demonstrate the effectiveness of our approach, which outperforms many recent related methods by a large margin. Code will be available athttps://github.com/ZhaofengSHI/CTDN. Zhaofeng Shi, Heqian Qiu, Lanxiao Wang, Fanman Meng, Qingbo Wu 0001, Hongliang Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Class Incremental Learning With Less Forgetting Direction and Equilibrium PointabstractCatastrophic forgetting is the core problem of class incremental learning (CIL). Existing work mainly adopts memory replay, knowledge distillation, and dynamic architecture to alleviate this problem, but seldom from the aspect of parameter regularization. However, existing parameter regularization methods struggle to achieve an appropriate balance between old and new tasks. To bring it back to CIL, we first propose constrained incremental learning with less forgetting direction (LFD) to leave more plasticity for the new task under a strong stability constraint for old tasks. Specifically, the new parameters are constrained to be close to the LFD of old tasks instead of a single group of old parameters. To validate the effectiveness of this regularization, we investigate the connectivity between the old parameters and the new parameters, and additionally find that a higher accuracy interval exists along the linear connection. Therefore, we further propose a post-processing procedure to find an equilibrium point in this interval for better balance between old and new tasks. Extensive classification experiments on CIFAR-100, ImageNet-100, and ImageNet-1K show our method can significantly improve performance compared with existing CIL methods and the object detection experiments on PASCAL-VOC show its broad generality on other tasks. Haitao Wen, Heqian Qiu, Lanxiao Wang, Haoyang Cheng, Hongliang Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Geodesic-Aligned Gradient Projection for Continual Task LearningabstractDeep networks notoriously suffer from performance deterioration on previous tasks when learning from sequential tasks, i.e., catastrophic forgetting. Recent methods of gradient projection show that the forgetting is resulted from the gradient interference on old tasks and accordingly propose to update the network in an orthogonal direction to the task space. However, these methods assume the task space is invariant and neglect the gradual change between tasks, resulting in sub-optimal gradient projection and a compromise of the continual learning capacity. To tackle this problem, we propose to embed each task subspace into a non-Euclidean manifold, which can naturally capture the change of tasks since the manifold is intrinsically non-static compared to the Euclidean space. Subsequently, we analytically derive the accumulated projection between any two subspaces on the manifold along the geodesic path by integrating an infinite number of intermediate subspaces. Building upon this derivation, we propose a novel geodesic-aligned gradient projection (GAGP) method that harnesses the accumulated projection to mitigate catastrophic forgetting. The proposed method utilizes the geometric structure information on the task manifold by capturing the gradual change between the new and the old tasks. Empirical studies on image classification demonstrate that the proposed method alleviates catastrophic forgetting and achieves on-par or better performance compared to the state-of-the-art approaches. Benliu Qiu, Heqian Qiu, Haitao Wen, Lanxiao Wang, Fanman Meng, Qingbo Wu 0001, Hongliang Li 0001 |
IEEE Trans. Image Process. | 4 |
| 2025 | Distribution-Level Memory Recall for Continual Learning: Preserving Knowledge and Avoiding ConfusionabstractContinual learning (CL) aims to enable deep neural networks (DNNs) to learn new data without forgetting previously learned knowledge. The key to achieving this goal is to avoid confusion at the feature level, i.e., to avoid confusion within old tasks and between new and old tasks. Existing prototype-based CL methods generate pseudo features for old knowledge replay by adding Gaussian noise to the centroids of old classes. However, the distribution in the feature space exhibits anisotropy during the incremental process, which prevents the pseudo features from faithfully reproducing the distribution of old knowledge in the feature space, leading to confusion at the classification boundaries within old tasks. To address this issue, we propose the distribution-level memory recall (DMR) method, which uses a Gaussian mixture model to precisely fit the feature distribution of old knowledge at the distribution level and generate pseudo features in the next stage. Furthermore, resistance to confusion at the distribution level is crucial for multimodal learning. Multimodal imbalance, which refers to uneven optimization processes among encoders of different modalities, results in significant differences in feature responses between modalities; this exacerbates confusion within old tasks in prototype-based CL methods. Therefore, we mitigate the multimodal imbalance problem by using the intermodal guidance and intramodal mining (IGIM) method to guide weaker modalities with prior information from dominant modalities and further explore useful information within modalities. To avoid confusion between new and old tasks, we propose using the confusion index to quantitatively describe a model's ability to distinguish between new and old tasks, and we use the incremental mixup feature enhancement (IMFE) method to enhance pseudo features with new sample features, alleviating classification confusion between new and old knowledge. We conduct extensive experiments on the CIFAR100, ImageNet100, TinyImageNet, ImageNet-1K and UESTC-MMEA-CL datasets and achieve state-of-the-art results. Shaoxu Cheng, Kanglei Geng, Chiyuan He, Zihuan Qiu, Linfeng Xu 0001, Heqian Qiu, Lanxiao Wang, Qingbo Wu 0001, Fanman Meng, Hongliang Li 0001 |
IEEE Trans. Multim. | 7 |
| 2024 | Prompt-Driven Referring Image Segmentation with Instance ContrastingabstractReferring image segmentation (RIS) aims to segment the target referent described by natural language. Recently, large-scale pre-trained models, e.g., CLIP and SAM, have been successfully applied in many downstream tasks, but they are not well adapted to RIS task due to inter-task differences. In this paper, we propose a new prompt-driven framework named Prompt-RIS, which bridges CLIP and SAM end-to-end and transfers their rich knowledge and powerful capabilities to RIS task through prompt learning. To adapt CLIP to pixel-level task, we first propose a Cross-Modal Prompting method, which acquires more comprehensive vision-language interaction and fine-grained text-to-pixel alignment by performing bidirectional prompting. Then, the prompt-tuned CLIP generates masks, points, and text prompts for SAM to generate more accurate mask predictions. Moreover, we further propose Instance Contrastive Learning to improve the model's discriminability to different instances and robustness to diverse languages describing the same instance. Extensive experiments demonstrate that the performance of our method outperforms the state-of-the-art methods consistently in both general and open-vocabulary settings. Chao Shang 0001, Zichen Song 0002, Heqian Qiu, Lanxiao Wang, Fanman Meng, Hongliang Li 0001 |
CVPR | 4 |
| 2024 | Class Incremental Learning with Multi-Teacher DistillationabstractDistillation strategies are currently the primary approaches for mitigating forgetting in class incremental learning (CIL). Existing methods generally inherit previous knowledge from a single teacher. However, teachers with different mechanisms are talented at different tasks, and inheriting diverse knowledge from them can enhance compatibility with new knowledge. In this paper, we propose the MTD method to find multiple diverse teachers for CIL. Specifically, we adopt weight permutation, feature perturbation, and diversity regularization techniques to ensure diverse mechanisms in teachers. To reduce time and memory consumption, each teacher is represented as a small branch in the model. We adapt existing CIL distillation strategies with MTD and extensive experiments on CIFAR-100, ImageNet-100, and ImageNet-1000 show significant performance improvement. Our code is available at https://github.com/HaitaoWen/CLearning. Haitao Wen, Lili Pan 0001, Heqian Qiu, Lanxiao Wang, Qingbo Wu 0001, Hongliang Li 0001 |
CVPR | 5 |
| 2024 | A Text Detector Based on the Specific Text PromptabstractNowadays, the prompt tuning has emerged as a novel new paradigm for adapting the original large-scale Contrastive Language-Image Pre-trained (CLIP) model into the downstream task as text detection. However, the learnable prompt adopted by the existing methods of prompt-tuning represents blurry and abstract meanings instead of fine-grained text feature. In this paper, we propose a powerful and robust text detector, called STP-TD, utilizing the specific text prompt and a learnable visual mask to fully apply the prior knowledge of CLIP model into the downstream task of text detection. STP-TD aims to make the split prompt character focusing on an ordered image token by Transformer mechanism. It is anticipated that a prompt character could stand for the fine-grained text feature of an image token, thus a distance loss is added to rectify prompt through optimization. Additionally, STP-TD firstly proposes a learnable visual mask to refine the text region in advance. Meanwhile, a synergetic framework is introduced as a bridge between visual branch and text branch. We also adopt the pixel-text matching process to align every pixel of image feature with the text feature. The experiments are conducted on the datasets ICDAR2015 and TotalText and outperform the state of the art. Xingtao Lin, Chuanyang Gong, Lanxiao Wang, Heqian Qiu, Shengyu Tong, Hongliang Li 0001 |
ICIP | 3 |
| 2024 | Video Class-Incremental Learning With Clip Based TransformerabstractVision Language Pre-training Models have shown significant potential in various domains, but there are few attempts to introduce it in the field of continual learning for video action recognition. We propose Video Class-Incremental Learner with CLIP based Transformer (VCIL-CT), which uses CLIP based vision transformer to train action recognition task by class-incremental learning pipeline. To specifically address the issue of catastrophic forgetting in transformer, we introduce Attention Distillation which distilling the attention feature from each transformer decoder. In the process of incremental learning of classes, there may be a problem of high bias towards new classes, we incorporate Class Balance Module to prevent bias on new task. Furthermore, we adopt Exemplar Augment strategy to improve exemplar quality on data replay step. We evaluate our proposed method based on the incremental action recognition benchmark presented by TCD, using UCF101, HMDB51, and UESTC-MMEA-CL datasets, and demonstrate the effectiveness of our algorithm compared to existing state-of-the-art continuous learning methods for action recognition. Shuyun Lu, Lanxiao Wang, Heqian Qiu, Xingtao Lin, Hefei Mei, Hongliang Li 0001 |
ICIP | 3 |
| 2024 | Attribute-Prompting Multi-Modal Object Reasoning Transformer for Remote Sensing Visual GroundingabstractRemote sensing visual grounding (RSVG) task aims to locate the particular object in a remote sensing image referred to a natural language expression, which requires to precisely fuse and align features from different modalities. However, existing methods usually use object-based multi-modal fusion, which is limited to capturing the detailed object characteristics in remote sensing images, resulting in object confusion with similar objects. To address this problem, we propose an attribute-prompting multi-modal object reasoning network for RSVG. Specifically, we first develop a learnable attribute prompter to adaptively explore diverse and rich attribute information according to common object characteristics in RS. With the help of attribute prompts, we design an attribute-prompting multi-modal fusion encoder to build fine-grained interactive and alignment between the visual and language features to avoid object confusion. Furthermore, we design a multi-modal progressive object reasoning decoder to gradually query more comprehensive object features for accurate object localization. Experimental results demonstrate that the proposed method achieves significant improvements. Heqian Qiu, Lanxiao Wang, Minjian Zhang 0003, Taijin Zhao, Hongliang Li 0001 |
IGARSS | 2 |
| 2024 | DP-RSCAP: Dual Prompt-Based Scene and Entity Network for Remote Sensing Image CaptioningabstractAs a challenging task towards remote sensing image analysis, the core problem of remote sensing image captioning is how to accurately transform the vision information into text information. Existing methods usually achieve it based on the simple multi-task learning strategy or visual attention mechanism, which ignores the importance of intermediate connection information for cross-modal transformation. To solve above problem, we propose a novel dual prompt-based scene and entity network (DP-RSCap) which aims to fully utilize the ability of cross-modal alignment in vision-language model build text prior information as intermediate connection to narrow the gap between different modalities and improve the quality of caption. Specifically, we first introduce an entity-concept prompt exporter to obtain explicit entity concepts in images. Then, we design a scene class prompt generator which can predict scene class and obtain fine-grained visual semantic features. Finally, we further design a dual prompt-based caption decoder to align and merge the visual semantic feature and dual prompts information as explicit intermediate connections, which can assist in generating precise caption. Extensive experiments on the challenging RSICD demonstrate the superior ability of our model. Lanxiao Wang, Heqian Qiu, Minjian Zhang 0003, Fanman Meng, Qingbo Wu 0001, Hongliang Li 0001 |
IGARSS | 1 |
| 2024 | Proposal-level Correction Guided by CLIP for Few-shot Object DetectionabstractFew-shot object detection aims at detecting previously unseen objects given only a few annotated samples. Most existing approaches treat the model obtained from the base training stage with abundant data as a container of prior knowledge that can be transferred to novel objects. Knowledge with similar properties is also contained in Contrastive Language-Image Pretraining (CLIP). In this paper, we utilize this external prior knowledge to generate proposal-level classification scores to improve the detection results. We notice that these scores can hardly reflect the quality of proposal localization, so we combine them with the ones from a conventional detector to obtain the ability to distinguish the background. Moreover, we propose a new score fusion module with regularization to alleviate the ambiguity of detection results generated by a trivial element-wise multiplication fusion method. To further improve the quality of classification scores in our proposed branch, we add learnable prompts to mitigate the inaccurate classification problem we observe. We conduct extensive experiments on the PASCAL VOC dataset and demonstrate the effectiveness of our approach. Ruihang Wang, Taijin Zhao, Hefei Mei, Heqian Qiu, Lanxiao Wang, Hongliang Li 0001 |
VCIP | 5 |
| 2024 | VLM-guided Explicit-Implicit Complementary novel class semantic learning for few-shot object detection
Taijin Zhao, Heqian Qiu, Lanxiao Wang, Hefei Mei, Fanman Meng, Qingbo Wu 0001, Hongliang Li 0001 |
Expert Syst. Appl. | 4 |
| 2024 | TridentCap: Image-Fact-Style Trident Semantic Framework for Stylized Image CaptioningabstractStylized image captioning (SIC) aims to generate captions with target style for images. The biggest challenge is that the collection and annotation of stylized data are pretty difficult and time-consuming. Most existing methods learn massive factual captions or additional stylized bookcorpus independently to assist in generating stylized caption, which ignore core relationships between existing image-fact-style trident data. In this paper, we propose a novel image-fact-style trident semantic framework TridentCap for stylized image captioning, which includes an image-fact semantic fusion encoder (SFE) and a trident stylization decoder (TSD). Unlike existing methods, we directly mine the core relationship in image-fact-style trident data and use factual semantic and image to build cross-modal semantic feature space, achieving the coherence between image and text. Specifically, SFE aims to learn the image-related prior language knowledge information from factual text and leverage fine-grained region-level semantic correlations of image and factual text to achieve cross-modal semantic information alignment and integration. TSD is designed to decouple the dual-source fused semantic feature based on the target style to achieve stylized caption generation. In addition, we design a pseudo labels filter (PLF) to obtain and expand massive image-fact-style trident data by building pseudo stylized annotations for all image-fact data in traditional caption datasets, which can further strengthen stylized caption learning. It is a generic algorithm to solve the problem of insufficient data and can be used into any existing stylized caption models. We conduct extensive experiments on SentiCap and FlickrStyle datasets, which achieve consistently improvement on almost all metrics. Our code will be released at: https://github.com/WangLanxiao/TridentCap_Code. Lanxiao Wang, Heqian Qiu, Benliu Qiu, Fanman Meng, Qingbo Wu 0001, Hongliang Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Oriented-DINO: Angle Decoupling Prediction and Consistency Optimizing for Oriented Detection TransformerabstractConsidering the arbitrary orientation of remote sensing objects, accurate angle prediction plays a crucial role in achieving precise oriented object detection (OOD) of aerial scenes. Existing transformer-based methods typically adopt an iterative refinement mechanism to update angle prediction and perform bipartite graph matching based on the combined matching costs. However, these methods may suffer from angle error accumulation across decoder layers and inconsistency between the L1 cost and the rotated intersection-of-union (IoU) cost, thus resulting in inaccurate angle prediction. To address these problems, this article proposes a novel transformer-based OOD method named Oriented-DINO (ODINO), which comprises three important components: error-mitigating angle decoupling prediction (EADP) module, nonlinear angle-conversion consistency optimizer (NACO), and query-driven diversity (QD) loss. To mitigate the angle error, the EADP module decouples angle prediction from the iterative box refinement process and uses independent branches to directly predict the angle. To address the issue of inconsistent matching, the NACO module uses a nonlinear function for angle conversion in matching cost calculation. This approach effectively alleviates the matching cost discrepancy in angle boundary case, while preserving the consistency in other instances. To avoid highly overlapped predictions triggered by similar queries, we introduce the QD loss to encourage the generation of diverse object queries, thus avoiding redundant predictions and enhancing prediction accuracy. Extensive experimental results demonstrate that our method achieves superior performance on OOD task. Minjian Zhang 0003, Heqian Qiu, Lanxiao Wang, Haoyang Cheng, Taijin Zhao, Hongliang Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | CrowdCaption++: Collective-Guided Crowd Scenes CaptioningabstractCrowd scenes analysis plays an important role in various fields, including public security, smart cities, and intelligent transportation systems. However, traditional crowd scenes captioning methods mainly focus on a single and prominent crowd collective, which limits their ability to describe the different crowd collectives in complex crowd scenes. To address this issue, we propose a collective-guided crowd scenes captioning model (CrowdCaption++) to explore a more comprehensive and detailed description. We design a crowd features encoder (CFE) including double-query features encoder and foreground crowd features encoder, which uses double-query attention module (DQ-ATT) to capture more representative visual features and extracts foreground crowd features to avoid interference from background for collectives prediction. Moreover, we build a collective-guided captioning decoder (CCD) to generate captions of different crowd collectives without requiring extra alignment between crowd collectives and captions. To achieve this, we first design a crowd collectives predictor to identify multiple potential crowd collectives and create crowd collectives guidance information. Finally, we use the crowd collectives guidance information to merge useful visual features and further generate corresponding caption. We evaluate our approach on the latest crowd scenes dataset CrowdCaption and demonstrate that our model can achieve a comprehensive understanding and describe the different crowd collectives in complex crowd scenes. Lanxiao Wang, Hongliang Li 0001, Minjian Zhang 0003, Heqian Qiu, Fanman Meng, Qingbo Wu 0001, Linfeng Xu 0001 |
IEEE Trans. Multim. | 1 |
| 2024 | Learning Offset Probability Distribution for Accurate Object DetectionabstractObject detection combines object classification and object localization problems. Current object detection methods heavily depend on regression networks to locate objects, which are optimized with various regression loss functions to predict offsets between candidate boxes and objects. However, these regression losses are difficult to assign the appropriate penalties for samples with large offset errors, resulting in suboptimal regression networks and inaccurate object offsets. In this article, we consider object location as offset bin classification problem, and propose a distance-aware offset bin classification network optimized with multiple binary cross entropy losses to learn various offset probability distribution, including single label distribution and distance-aware label distribution. On one hand, it provides gradient contributions for different samples based on the bounded probability instead of previous incalculable offset error. On the other hand, it explores the distance correlations between discrete offset bins to facilitate network learning. Specifically, we discretize the continuous offset into a number of bins, and predict the probability of each offset bin, in which the probability should be higher for the offset bin closer to the target offsets, and vice versa. Furthermore, we propose an expectation-based offset prediction and a hierarchical focusing method to improve the precision of prediction. We conduct extensive experiments to evaluate the effectiveness of our method. In addition, our method can be conveniently and flexibly inserted into existing object detection methods, which consistently achieves a large gain based on popular anchor-based and anchor-free methods on the PASCAL VOC, MS-COCO, KITTI, and CrowdHuman datasets. Code will be released at: https://github.com/QiuHeqian/DBC . Heqian Qiu, Hongliang Li 0001, Qingbo Wu 0001, Hengcan Shi, Lanxiao Wang, Fanman Meng, Linfeng Xu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2023 | CafeBoost: Causal Feature Boost to Eliminate Task-Induced Bias for Class Incremental LearningabstractContinual learning requires a model to incrementally learn a sequence of tasks and aims to predict well on all the learned tasks so far, which notoriously suffers from the catastrophic forgetting problem. In this paper, we find a new type of bias appearing in continual learning, coined as task-induced bias. We place continual learning into a causal framework, based on which we find the task-induced bias is reduced naturally by two underlying mechanisms in task and domain incremental learning. However, these mechanisms do not exist in class incremental learning (CIL), in which each task contains a unique subset of classes. To eliminate the task-induced bias in CIL, we devise a causal intervention operation so as to cut off the causal path that causes the task-induced bias, and then implement it as a causal debias module that transforms biased features into unbiased ones. In addition, we propose a training pipeline to incorporate the novel module into existing methods and jointly optimize the entire architecture. Our overall approach does not rely on data replay, and is simple and convenient to plug into existing methods. Extensive empirical study on CIFAR-100 and ImageNet shows that our approach can improve accuracy and reduce forgetting of well-established methods by a large margin. Benliu Qiu, Hongliang Li 0001, Haitao Wen, Heqian Qiu, Lanxiao Wang, Fanman Meng, Qingbo Wu 0001, Lili Pan 0001 |
CVPR | 5 |
| 2023 | Incrementer: Transformer for Class-Incremental Semantic Segmentation with Knowledge Distillation Focusing on Old ClassabstractClass-incremental semantic segmentation aims to incrementally learn new classes while maintaining the capability to segment old ones, and suffers catastrophic forgetting since the old-class labels are unavailable. Most existing methods are based on convolutional networks and prevent forgetting through knowledge distillation, which (1) need to add additional convolutional layers to predict new classes, and (2) ignore to distinguish different regions corresponding to old and new classes during knowledge distillation and roughly distill all the features, thus limiting the learning of new classes. Based on the above observations, we propose a new transformer framework for class-incremental semantic segmentation, dubbed Incrementer, which only needs to add new class tokens to the transformer decoder for new-class learning. Based on the Incrementer, we propose a new knowledge distillation scheme that focuses on the distillation in the old-class regions, which reduces the constraints of the old model on the new-class learning, thus improving the plasticity. Moreover, we propose a class deconfusion strategy to alleviate the overfitting to new classes and the confusion of similar classes. Our method is simple and effective, and extensive experiments show that our method outperforms the SOTAs by a large margin (5~15 absolute points boosts on both Pascal VOC and ADE20k). We hope that our Incrementer can serve as a new strong pipeline for class-incremental semantic segmentation. Chao Shang 0001, Hongliang Li 0001, Fanman Meng, Qingbo Wu 0001, Heqian Qiu, Lanxiao Wang |
CVPR | 6 |
| 2023 | Contrastive Continuity on Augmentation Stability Rehearsal for Continual Self-Supervised LearningabstractSelf-supervised learning has attracted a lot of attention recently, which is able to learn powerful representations without any manual annotations. However, self-supervised learning needs to develop the ability to continuously learn to cope with a variety of real-world challenges, i.e., Continual Self-Supervised Learning (CSSL). Catastrophic forgetting is a notorious problem in CSSL, where the model tends to forget the learned knowledge. In practice, simple rehearsal or regularization will bring extra negative effects while alleviating catastrophic forgetting in CSSL, e.g., overfitting on the rehearsal samples or hindering the model from encoding fresh information. In order to address catastrophic forgetting without overfitting on the rehearsal samples, we propose Augmentation Stability Rehearsal (ASR) in this paper, which selects the most representative and discriminative samples by estimating the augmentation stability for rehearsal. Meanwhile, we design a matching strategy for ASR to dynamically update the rehearsal buffer. In addition, we further propose Contrastive Continuity on Augmentation Stability Rehearsal (C2ASR) based on ASR. We show that C2ASR is an upper bound of the Information Bottleneck (IB) principle, which suggests that C2ASR essentially preserves as much information shared among seen task streams as possible to prevent catastrophic forgetting and dismisses the redundant information between previous task streams and current task stream to free up the ability to encode fresh information. Our method obtains a great achievement compared with state-of-the-art CSSL methods on a variety of CSSL benchmarks. Haoyang Cheng, Haitao Wen, Xiaoliang Zhang 0002, Heqian Qiu, Lanxiao Wang, Hongliang Li 0001 |
ICCV | 5 |
| 2023 | Optimizing Mode Connectivity for Class Incremental LearningabstractClass incremental learning (CIL) is one of the most challenging scenarios in continual learning. Existing work mainly focuses on strategies like memory replay, regularization, or dynamic architecture but ignores a crucial aspect: mode connectivity. Recent studies have shown that different minima can be connected by a low-loss valley, and ensembling over the valley shows improved performance and robustness. Motivated by this, we try to investigate the connectivity in CIL and find that the high-loss ridge exists along the linear connection between two adjacent continual minima. To dodge the ridge, we propose parameter-saving OPtimizing Connectivity (OPC) based on Fourier series and gradient projection for finding the low-loss path between minima. The optimized path provides infinite low-loss solutions. We further propose EOPC to ensemble points within a local bent cylinder to improve performance on learned tasks. Our scheme can serve as a plug-in unit, extensive experiments on CIFAR-100, ImageNet-100, and ImageNet-1K show consistent improvements when adapting EOPC to existing representative CIL methods. Our code is available at https://github.com/HaitaoWen/EOPC. Haitao Wen, Haoyang Cheng, Heqian Qiu, Lanxiao Wang, Lili Pan 0001, Hongliang Li 0001 |
ICML | 4 |
| 2023 | CrossDet++: Growing Crossline Representation for Object DetectionabstractIn object detection, precise object representation is a key factor to successfully classify and locate objects of an image. Existing methods usually use rectangular anchor boxes or a set of points to represent objects. However, these methods either introduce background noise or miss the continuous appearance information inside the object, and thus cause incorrect detection results. In this paper, we propose a novel anchor-free object detection network, called CrossDet++, which uses a set of growing crosslines along horizontal and vertical axes as object representations. An object can be flexibly represented as crosslines in different combinations, which inspires us to select the expressive crossline to effectively reduce the interference of noise. Meanwhile, the crossline representation takes into account the continuous adjacent object information, which is useful to enhance the discriminability of object features and find the object boundaries. Based on the learned crosslines, we propose an axis-query crossline growing module to adaptively capture features of crosslines and query surrounding pixels related to the line features for subsequent growing of crosslines. Their growing offsets and scales can be supervised by a decoupled regression mechanism, which limits the regression target to a specific direction for decreasing the optimization difficulty. During the training, we design a semantic-guided label assignment to emphasize the importance of crossline targets with higher semantic richness, further improving the detection performance. The experiment results demonstrate the effectiveness of our proposed method. Code can be available at:https://github.com/QiuHeqian/CrossDet. Heqian Qiu, Hongliang Li 0001, Qingbo Wu 0001, Jianhua Cui, Zichen Song 0002, Lanxiao Wang, Minjian Zhang 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2023 | DRDet: Dual-Angle Rotated Line Representation for Oriented Object DetectionabstractIn aerial scenes, oriented object detection is sensitive to the orientation of objects, which makes the formulation of orientation-aware object representation become a critical problem. Existing methods mostly adopt rectangle anchor or discrete points as object representation, which may lead to the feature aliasing between overlapping objects and ignore the orientation information of objects. To solve these issues, we propose a novel anchor-free oriented object detection network named DRDet, which adopts Dual-angle Rotated Lines (DRL) as object representation. Different from other object representations, DRL can adaptively rotate and extend to the boundary of the object according to its orientation and shape, which explicitly introduces the orientation information into the formulation of object representation. And it can adaptively cope with the geometric deformation of objects. Based on the dual-angle rotated lines, we design an Orientation-guided Feature Encoder (OFE) to encode discriminant object feature along each rotated line, respectively. Instead of encoding rectangle feature, the OFE module adopts line features for orientation-guided feature encoding, which can alleviate the feature aliasing between neighboring objects or background. To further enhance the flexibility of dual-angle rotated lines, we design a Dual-angle Decoder (DD) that predicts two angle offsets according to the orientation-guided feature and converts the angle offsets and regression offsets into dual-angle rotated line representation, which can help to guide the adaptive rotation of each rotated line, respectively. Our proposed method achieves consistent improvement on both DOTA and HRSC2016 datasets. Extensive experimental results verify the effectiveness of our method in oriented object detection. Minjian Zhang 0003, Heqian Qiu, Hefei Mei, Lanxiao Wang, Fanman Meng, Linfeng Xu 0001, Hongliang Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | What Happens in Crowd Scenes: A New Dataset About Crowd Scenes for Image CaptioningabstractMaking machines endowed with eyes and brains to effectively understand and analyze crowd scenes is of paramount importance for building a smart city to serve people. This is of far-reaching significance for the guidance of dense crowds and accident prevention, such as crowding and stampedes. As a typical multimodal scene understanding task, image captioning has always attracted widespread attention. However, crowd scene understanding captioning is rarely studied due to the unobtainability of related datasets. Therefore, it is difficult to know what happens in crowd scenes. In order to fill this research gap, we propose a crowd scenes caption dataset named CrowdCaption which has the advantages of crowd-topic scenes, comprehensive and complex caption descriptions, typical relationships and detailed grounding annotations. The complexity and diversity of the descriptions and the specificity of the crowd scenes make this dataset extremely challenging to most current methods. Thus, we propose a Multi-hierarchical Attribute Guided Crowd Caption Network (MAGC) based on crowd objects, actions, and status (such as position, dress, posture, etc.) aiming to generate crowd-specific detailed descriptions. We conduct extensive experiments on our CrowdCaption dataset, and our proposed method reaches the state-of-the-art (SoTA) performance. We hope the CrowdCaption dataset can assist future studies related to crowd scenes in the multimodal domain. Lanxiao Wang, Hongliang Li 0001, Wenzhe Hu, Xiaoliang Zhang 0002, Heqian Qiu, Fanman Meng, Qingbo Wu 0001 |
IEEE Trans. Multim. | 1 |
| 2022 | Spatial-Semantic Attention for Grounded Image CaptioningabstractGrounded image captioning models usually process high-dimensional vectors from the feature extractor to generate descriptions. However, mere vectors do not provide adequate information. The model needs more explicit information for grounded image captioning. Besides high dimensional vectors, the feature extractor also predicts the locations and categories of the objects, which contains low-level spatial information and high-level semantic information. To this end, we propose a new attention module called Spatial-Semantic (SS) Attention, which utilizes the predictions from the backbone network to help the model attend to the correct objects. Specifically, the SS attention module collects the position of proposals and the class probabilities from the feature extractor as spatial and semantic information to assist attention weighting. In addition, we propose a grounding loss to supervise the SS attention. Our method achieves high performance on captioning and grounding metrics and outperforms some powerful previous models on the Flickr30k Entities dataset. Wenzhe Hu, Lanxiao Wang, Linfeng Xu 0001 |
ICIP | 2 |
| 2022 | RefCrowd: Grounding the Target in Crowd with Referring ExpressionsabstractCrowd understanding has aroused the widespread interest in vision domain due to its important practical significance. Unfortunately, there is no effort to explore crowd understanding in multi-modal domain that bridges natural language and computer vision. Referring expression comprehension (REF) is such a representative multi-modal task. Current REF studies focus more on grounding the target object from multiple distinctive categories in general scenarios. It is difficult to applied to complex real-world crowd understanding. To fill this gap, we propose a new challenging dataset, called RefCrowd, which towards looking for the target person in crowd with referring expressions. It not only requires to sufficiently mine natural language information, but also requires to carefully focus on subtle differences between the target and a crowd of persons with similar appearance, so as to realize fine-grained mapping from language to vision. Furthermore, we propose a Fine-grained Multi-modal Attribute Contrastive Network (FMAC) to deal with REF in crowd understanding. It first decomposes the intricate visual and language features into attribute-aware multi-modal features, and then captures discriminative but robustness fine-grained attribute features to effectively distinguish these subtle differences between similar persons. The proposed method outperforms existing state-of-the-art (SoTA) methods on our RefCrowd dataset and existing REF datasets. In addition, we implement an end-to-end REF toolbox for the deeper research in multi-modal domain. Our dataset and code can be available at: https://qiuheqian.github.io/datasets/refcrowd/. Heqian Qiu, Hongliang Li 0001, Taijin Zhao, Lanxiao Wang, Qingbo Wu 0001, Fanman Meng |
ACM Multimedia | 4 |
| 2022 | STSI: Efficiently Mine Spatio- Temporal Semantic Information between Different Multimodal for Video CaptioningabstractAs one of the challenging tasks in computer vision, video captioning needs to use natural language to describe the content of video. Video contains complex information, such as semantic information, time information and so on. How to synthesize sentences effectively from rich and different kinds of information is very significant. The existing methods often cannot well integrate the multimodal feature to predict the association between different objects in video. In this paper, we improve the existing encoder-decoder structure and propose a network deeply mining the spatio-temporal correlation between multimodal features. Through the analysis of sentence components, we use spatio-temporal semantic information mining module to fuse the object, 2D and 3D features in both time and space. It is worth mentioning that the word output at the previous time is added as the prediction branch of auxiliary conjunctions. After that, a dynamic gumbel scorer is used to output caption sentences that are more consistent with the facts. The experimental results on two benchmark datasets show that our STSI is superior to the state-of-the-art methods while generating more reasonable and semantic-logical sentences. Huiyu Xiong, Lanxiao Wang |
VCIP | 2 |
| 2022 | Real-time panoptic segmentation with relationship between adjacent pixels and boundary prediction
Xiaoliang Zhang 0002, Hongliang Li 0001, Lanxiao Wang, Haoyang Cheng, Heqian Qiu, Wenzhe Hu, Fanman Meng, Qingbo Wu 0001 |
Neurocomputing | 3 |
| 2022 | POS-Trends Dynamic-Aware Model for Video CaptionabstractVideo caption aims to generate descriptive sentences about the video, and the most critical problem is how to achieve accurate word prediction with standardized and coherent syntax structure, which requires the model to thoroughly understand video content and precisely map them into corresponding sentence components. Many existing methods usually fuse different video features into a single visual feature for generating sentences. However, they ignore the word dataset prior information in the annotations (such as Part-Of-Speech) and they also ignore the association between sentence components and types of visual features. To solve these problems, we propose a POS-trends dynamic-aware model (PDA) to fully exploit the word dataset prior information in the captions to predict POS tag, so as to assist generating captions. We propose a POS feature extraction (PFE) module to use different filters to extract different POS-trends features, predict POS tags and fuse visual features. Furthermore, we propose a visual-dynamic-aware (VDA) module to dynamically adjust the mapping way of words and supplement the visual information into the local features. The fusion features provide directional visual information to generate correct words, and the predicted POS tags to guide the decoding process to generate a more standardized and coherent syntax structure. A large number of experiments based on MSVD, MSR-VTT and VATEX demonstrated that our method outperforms the state-of-the-art methods in BLEU-4, ROUGE-L, METEOR, CIDEr. Code can be available at:https://github.com/WangLanxiao/PDA-for-video-caption. Lanxiao Wang, Hongliang Li 0001, Heqian Qiu, Qingbo Wu 0001, Fanman Meng, King Ngi Ngan |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2021 | CrossDet: Crossline Representation for Object DetectionabstractObject detection aims to accurately locate and classify objects in an image, which requires precise object representations. Existing methods usually use rectangular anchor boxes or a set of points to represent objects. However, these methods either introduce background noise or miss the continuous appearance information inside the object, and thus cause incorrect detection results. In this paper, we propose a novel anchor-free object detection network, called Cross-Det, which uses a set of growing cross lines along horizontal and vertical axes as object representations. An object can be flexibly represented as cross lines in different combinations. It not only can effectively reduce the interference of noise, but also take into account the continuous object information, which is useful to enhance the discriminability of object features and find the object boundaries. Based on the learned cross lines, we propose a crossline extraction module to adaptively capture features of cross lines. Furthermore, we design a decoupled regression mechanism to regress the localization along the horizontal and vertical directions respectively, which helps to decrease the optimization difficulty because the optimization space is limited to a specific direction. Our method achieves consistently improvement on the PASCAL VOC and MS-COCO datasets. The experiment results demonstrate the effectiveness of our proposed method. Code can be available at: https://github.com/QiuHeqian/CrossDet. Heqian Qiu, Hongliang Li 0001, Qingbo Wu 0001, Jianhua Cui, Zichen Song 0002, Lanxiao Wang, Minjian Zhang 0003 |
ICCV | 6 |
| 2020 | Multi-stage Tag Guidance Network in Video CaptionabstractRecently, video caption plays an important role in computer vision tasks. We participate in Pre-training for Video Captioning Challenge which aims to produce at least one sentence for each challenge video based on the pretraining models. In this work, we propose a tag guidance module to learn a representation which can better build the interaction in cross-modal between visual content and textual sentences. First, we utilize three types of features extraction networks to fully capture the information of 2D, 3D and object information. Second, to prevent overfitting and time issues, the entire process of training is divided into two stages. The first stage trains all data, and the second stage introduces a random dropout. Furthermore, we train a CNN-based network to pick out the best candidate results. In summary, we were ranked third place in Pre-training for Video Captioning Challenge which proved the effectiveness of our model. Lanxiao Wang, Chao Shang 0001, Heqian Qiu, Taijin Zhao, Benliu Qiu, Hongliang Li 0001 |
ACM Multimedia | 1 |