VLDB 2026 Research / reviewers in the wild / expert
Taijin Zhao
dblp:276/3258
· DBLP profile ↗
16ranked-venue papers
2as first author
13since 2021 · last 2026
0009-0008-5449-900XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Parameter Merging with Gradient-Guided Supermasks in Online Continual LearningabstractOnline continual learning (OCL) aims at learning a non-stationary data stream in a way of reading each data sample only once, and hence suffers from the trade-off of catastrophic forgetting and insufficient learning. In this work, we firstly analytically establish relationship between loss functions and model parameters from the Bayesian perspective. Based on our analysis, we subsequently propose a parameter merging method with gradient-guided supermasks. Our method leverages 1-order and 2-order gradient information to construct supermasks that determine the merging weights between the old and new models. Our method performs direct arithmetic operations on parameters to update models, beyond traditional gradient descent. We further discover that a widely-used premise that 1-order gradients can be negligible is invalid in OCL, due to slow convergence incurred by insufficient learning. Additionally, we utilize a dual-model dual-view distillation strategy that can align output distributions of the new and merged models for each sample, further enhancing model performance. Extensive experiments are conducted on four benchmarks in OCL settings, including CIFAR-10, CIFAR-100, Tiny-ImageNet, and ImageNet-100. Experimental results demonstrate that our method is effective, and achieves a substantial boost over previous methods. Benliu Qiu, Heqian Qiu, Lanxiao Wang, Taijin Zhao, Lili Pan 0001, Hongliang Li 0001 |
AAAI | 4 |
| 2026 | Ego-PMOVE: Prompt-aware Mixture of View Experts Network for Egocentric Gaze PredictionabstractEgocentric gaze prediction serves as a critical indicator for decoding human visual attention and cognitive processes, but its inherently limited field of view creates prediction challenges. Although exo-view data provides supplementary contextual information, it exhibits significant spatial and semantic gaps. Existing methods focus solely on isolated feature encoding in single-view paradigms, neglecting cross-view gaze correlations. To make up for this gap, we make the first exploration of cross-view gaze relationship for egocentric gaze prediction, and propose Ego-PMOVE, a novel Prompt-aware Mixture of View Experts network. Unlike prior cross-view studies that forcibly align cross-view features thereby introducing inference noise, we leverage the popular Mixture-of-Experts (MoE) and a set of flexible prompts to disentangle features from different views into three parallel experts: a view-shared expert directly modeling common semantic relationships, a view-discrepancy expert adaptively adjusting the spatial position, scale and shifts based on different view-specific features, and an egocentric expert extracting independent features to compensate for the case of missing exocentric data. To balance these experts, we further design a soft router to dynamically weight them for mining useful information while suppressing noise. A view-query gaze decoder then generates view-specific gaze attention maps, jointly optimized by gaze-heamap and cross-view contrastive loss that regularize both shared and divergent features for accurate gaze prediction. Extensive experiments across the multi-view EgoMe dataset and single-view Ego4D and EGTEA Gaze++ datasets demonstrate the effectiveness and generalizability of our approach. Heqian Qiu, Lanxiao Wang, Taijin Zhao, Zhaofeng Shi, Linfeng Xu 0001, Hongliang Li 0001 |
AAAI | 3 |
| 2025 | GRSDet: Learning to Generate Local Reverse Samples for Few-shot Object Detection
Hefei Mei, Taijin Zhao, Shiyuan Tang, Heqian Qiu, Lanxiao Wang, Minjian Zhang 0003, Fanman Meng, Hongliang Li 0001 |
Neurocomputing | 2 |
| 2025 | Adaptively forget with crossmodal and textual distillation for class-incremental video captioning
Huiyu Xiong, Lanxiao Wang, Heqian Qiu, Taijin Zhao, Benliu Qiu, Hongliang Li 0001 |
Neurocomputing | 4 |
| 2025 | MCCE-REC: MLLM-Driven Cross-Modal Contrastive Entropy Model for Zero-Shot Referring Expression ComprehensionabstractZero-shot referring expression comprehension (zero-shot REC) is a crucial yet challenging task in the field of multi-modal understanding, which aims to locate an object described by a referring expression without training on task-specific datasets. Existing methods take advantage of a pre-trained CLIP model to align cropped proposal regions with referring expressions. However, our analysis reveals that this aligning way heavily biases toward certain salient visual regions due to CLIP focusing on global-level image-text matching. To mitigate this bias, we propose MCCE-REC, an MLLM-driven cross-modal contrastive entropy model for training-free zero-shot REC. Benefiting from the remarkable in-context comprehension ability of the multi-modal large language model (MLLM), we design a set of referring prompts for MLLM to generate diverse detailed informative, and contrastive cues related to referring objects. Based on these cues, on the one hand, we propose a multi-cues cross-modal interaction network, which associates the visual features and referring object textual features from multiple perspectives and perceives surrounding context object information in a parameter-free manner, avoiding bias towards salient features. On the other hand, we introduce a contrastive similarity entropy selection mechanism that compares the positive and negative cues to suppress biased regions with high similarity scores and emphasizes accurate regions correlating with referring descriptions. Extensive experiments demonstrate our MCCE-REC outperforms existing zero-shot methods by a significant margin on various REC datasets. Heqian Qiu, Lanxiao Wang, Taijin Zhao, Fanman Meng, Qingbo Wu 0001, Hongliang Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Attribute-Prompting Multi-Modal Object Reasoning Transformer for Remote Sensing Visual GroundingabstractRemote sensing visual grounding (RSVG) task aims to locate the particular object in a remote sensing image referred to a natural language expression, which requires to precisely fuse and align features from different modalities. However, existing methods usually use object-based multi-modal fusion, which is limited to capturing the detailed object characteristics in remote sensing images, resulting in object confusion with similar objects. To address this problem, we propose an attribute-prompting multi-modal object reasoning network for RSVG. Specifically, we first develop a learnable attribute prompter to adaptively explore diverse and rich attribute information according to common object characteristics in RS. With the help of attribute prompts, we design an attribute-prompting multi-modal fusion encoder to build fine-grained interactive and alignment between the visual and language features to avoid object confusion. Furthermore, we design a multi-modal progressive object reasoning decoder to gradually query more comprehensive object features for accurate object localization. Experimental results demonstrate that the proposed method achieves significant improvements. Heqian Qiu, Lanxiao Wang, Minjian Zhang 0003, Taijin Zhao, Hongliang Li 0001 |
IGARSS | 4 |
| 2024 | Proposal-level Correction Guided by CLIP for Few-shot Object DetectionabstractFew-shot object detection aims at detecting previously unseen objects given only a few annotated samples. Most existing approaches treat the model obtained from the base training stage with abundant data as a container of prior knowledge that can be transferred to novel objects. Knowledge with similar properties is also contained in Contrastive Language-Image Pretraining (CLIP). In this paper, we utilize this external prior knowledge to generate proposal-level classification scores to improve the detection results. We notice that these scores can hardly reflect the quality of proposal localization, so we combine them with the ones from a conventional detector to obtain the ability to distinguish the background. Moreover, we propose a new score fusion module with regularization to alleviate the ambiguity of detection results generated by a trivial element-wise multiplication fusion method. To further improve the quality of classification scores in our proposed branch, we add learnable prompts to mitigate the inaccurate classification problem we observe. We conduct extensive experiments on the PASCAL VOC dataset and demonstrate the effectiveness of our approach. Ruihang Wang, Taijin Zhao, Hefei Mei, Heqian Qiu, Lanxiao Wang, Hongliang Li 0001 |
VCIP | 2 |
| 2024 | VLM-guided Explicit-Implicit Complementary novel class semantic learning for few-shot object detection
Taijin Zhao, Heqian Qiu, Lanxiao Wang, Hefei Mei, Fanman Meng, Qingbo Wu 0001, Hongliang Li 0001 |
Expert Syst. Appl. | 1 |
| 2024 | Oriented-DINO: Angle Decoupling Prediction and Consistency Optimizing for Oriented Detection TransformerabstractConsidering the arbitrary orientation of remote sensing objects, accurate angle prediction plays a crucial role in achieving precise oriented object detection (OOD) of aerial scenes. Existing transformer-based methods typically adopt an iterative refinement mechanism to update angle prediction and perform bipartite graph matching based on the combined matching costs. However, these methods may suffer from angle error accumulation across decoder layers and inconsistency between the L1 cost and the rotated intersection-of-union (IoU) cost, thus resulting in inaccurate angle prediction. To address these problems, this article proposes a novel transformer-based OOD method named Oriented-DINO (ODINO), which comprises three important components: error-mitigating angle decoupling prediction (EADP) module, nonlinear angle-conversion consistency optimizer (NACO), and query-driven diversity (QD) loss. To mitigate the angle error, the EADP module decouples angle prediction from the iterative box refinement process and uses independent branches to directly predict the angle. To address the issue of inconsistent matching, the NACO module uses a nonlinear function for angle conversion in matching cost calculation. This approach effectively alleviates the matching cost discrepancy in angle boundary case, while preserving the consistency in other instances. To avoid highly overlapped predictions triggered by similar queries, we introduce the QD loss to encourage the generation of diverse object queries, thus avoiding redundant predictions and enhancing prediction accuracy. Extensive experimental results demonstrate that our method achieves superior performance on OOD task. Minjian Zhang 0003, Heqian Qiu, Lanxiao Wang, Haoyang Cheng, Taijin Zhao, Hongliang Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | Novel-Registrable Weights and Region-Level Contrastive Learning for Incremental Few-shot Object Detection
Shiyuan Tang, Hefei Mei, Heqian Qiu, Xinpeng Hao, Taijin Zhao, Benliu Qiu, Haoyang Cheng, Chuanyang Gong, Hongliang Li 0001 |
ICONIP (11) | 5 |
| 2023 | Disturbed Augmentation Invariance for Unsupervised Visual Representation LearningabstractContrastive learning has gained great prominence recently, which achieves excellent performance by simple augmentation invariance. However, the simple contrastive pairs suffer from lacking of diversity due to the mechanical augmentation strategies. In this paper, we propose Disturbed Augmentation Invariance (DAI for abbreviation), which constructs disturbed contrastive pairs by generating appropriate disturbed views for each augmented view in the feature space to increase the diversity. In practice, we establish a multivariate normal distribution for each augmented view, whose mean is corresponding augmented view and covariance matrix is estimated from its nearest neighbors in the dataset. Then we sample random vectors from this distribution as the disturbed views to construct disturbed contrastive pairs. In order to avoid extra computational cost with the increase of disturbed contrastive pairs, we utilize an upper bound of the trivial disturbed augmentation invariance loss to construct the DAI loss. In addition, we propose Bottleneck version of Disturbed Augmentation Invariance (BDAI for abbreviation) inspired by the Information Bottleneck principle, which further refines the extracted information and learns a compact representation by additionally increasing the variance of the original contrastive pair. In order to make BDAI work effectively, we design a statistical strategy to control the balance between the amount of the information shared by all disturbed contrastive pairs and the compactness of the representation. Our approach gets a consistent improvement over the popular contrastive learning methods on a variety of downstream tasks, e.g. image classification, object detection and instance segmentation. Haoyang Cheng, Hongliang Li 0001, Qingbo Wu 0001, Heqian Qiu, Xiaoliang Zhang 0002, Fanman Meng, Taijin Zhao |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2023 | Cross-Modal Recurrent Semantic Comprehension for Referring Image SegmentationabstractReferring image segmentation aims to segment the target object from the image according to the description of language expression. Due to the diversity of language expressions, word sequences in different orders often express different semantic information. The previous methods focus more on matching different words to different visual regions in the image separately, ignoring the global semantic understanding of language expression based on the sequence structure. To address this problem, we redesign a new recurrent network structure for referring image segmentation, called Cross-Modal Recurrent Semantic Comprehension Network (CRSCNet), to obtain a more comprehensive global semantic understanding through iterative cross-modal semantic reasoning. Specifically, in each iteration, we first propose a Dynamic SepConv to extract relevant visual features guided by language and further propose Language Attentional Feature Modulation to improve the feature discriminability, then propose a Cross-Modal Semantic Reasoning module to perform global semantic reasoning by capturing both linguistic and visual information, and finally updates and corrects the visual features of the predicted object based on semantic information. Moreover, we further propose a Cross-Modal ASPP to capture richer visual information referred to in the global semantics of the language expression from larger receptive fields. Extensive experiments demonstrate that our proposed network significantly outperforms previous state-of-the-art methods on multiple datasets. Chao Shang 0001, Hongliang Li 0001, Heqian Qiu, Qingbo Wu 0001, Fanman Meng, Taijin Zhao, King Ngi Ngan |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2022 | RefCrowd: Grounding the Target in Crowd with Referring ExpressionsabstractCrowd understanding has aroused the widespread interest in vision domain due to its important practical significance. Unfortunately, there is no effort to explore crowd understanding in multi-modal domain that bridges natural language and computer vision. Referring expression comprehension (REF) is such a representative multi-modal task. Current REF studies focus more on grounding the target object from multiple distinctive categories in general scenarios. It is difficult to applied to complex real-world crowd understanding. To fill this gap, we propose a new challenging dataset, called RefCrowd, which towards looking for the target person in crowd with referring expressions. It not only requires to sufficiently mine natural language information, but also requires to carefully focus on subtle differences between the target and a crowd of persons with similar appearance, so as to realize fine-grained mapping from language to vision. Furthermore, we propose a Fine-grained Multi-modal Attribute Contrastive Network (FMAC) to deal with REF in crowd understanding. It first decomposes the intricate visual and language features into attribute-aware multi-modal features, and then captures discriminative but robustness fine-grained attribute features to effectively distinguish these subtle differences between similar persons. The proposed method outperforms existing state-of-the-art (SoTA) methods on our RefCrowd dataset and existing REF datasets. In addition, we implement an end-to-end REF toolbox for the deeper research in multi-modal domain. Our dataset and code can be available at: https://qiuheqian.github.io/datasets/refcrowd/. Heqian Qiu, Hongliang Li 0001, Taijin Zhao, Lanxiao Wang, Qingbo Wu 0001, Fanman Meng |
ACM Multimedia | 3 |
| 2020 | Language-Aware Fine-Grained Object Representation for Referring Expression ComprehensionabstractReferring expression comprehension expects to accurately locate an object described by a language expression, which requires precise language-aware visual object representations. However, existing methods usually use rectangular object representations, such as object proposal regions and grid regions. They ignore some fine-grained object information like shapes and poses, which are often described in language expressions and important to localize objects. Additionally, rectangular object regions usually contain background contents and irrelevant foreground features, which also decrease the localization performance. To address these problems, we propose a language-aware deformable convolution model (LDC) to learn language-aware fine-grained object representations. Rather than extracting rectangular object representations, LDC adaptively samples a set of key points based on the image and language to represent objects. This type of object representations can capture more fine-grained object information (e.g., shapes and poses) and suppress noises in accordance with language and thus, boosts the object localization performance. Based on the language-aware fine-grained object representation, we next design a bidirectional interaction model (BIM) that leverages a modified co-attention mechanism to build cross-modal bidirectional interactions to further improve the language and object representations. Furthermore, we propose a hierarchical fine-grained representation network (HFRN) to learn language-aware fine-grained object representations and cross-modal bidirectional interactions at local word level and global sentence level, respectively. Our proposed method outperforms the state-of-the-art methods on the RefCOCO, RefCOCO+ and RefCOCOg datasets. Heqian Qiu, Hongliang Li 0001, Qingbo Wu 0001, Fanman Meng, Hengcan Shi, Taijin Zhao, King Ngi Ngan |
ACM Multimedia | 6 |
| 2020 | Multi-stage Tag Guidance Network in Video CaptionabstractRecently, video caption plays an important role in computer vision tasks. We participate in Pre-training for Video Captioning Challenge which aims to produce at least one sentence for each challenge video based on the pretraining models. In this work, we propose a tag guidance module to learn a representation which can better build the interaction in cross-modal between visual content and textual sentences. First, we utilize three types of features extraction networks to fully capture the information of 2D, 3D and object information. Second, to prevent overfitting and time issues, the entire process of training is divided into two stages. The first stage trains all data, and the second stage introduces a random dropout. Furthermore, we train a CNN-based network to pick out the best candidate results. In summary, we were ranked third place in Pre-training for Video Captioning Challenge which proved the effectiveness of our model. Lanxiao Wang, Chao Shang 0001, Heqian Qiu, Taijin Zhao, Benliu Qiu, Hongliang Li 0001 |
ACM Multimedia | 4 |
| 2020 | A multi-scale language embedding network for proposal-free referring expression comprehensionabstractReferring expression comprehension (REC) is a task that aims to find the location of an object specified by a language expression. Current solutions for REC can be classified into proposal-based methods and proposal-free methods. Proposal-free methods are popular recently because of its flexibility and lightness. Nevertheless, existing proposal-free works give little consideration to visual context. As REC is a context sensitive task, it is hard for current proposal-free methods to comprehend expressions that describe objects by the relative position with surrounding things. In this paper, we propose a multi-scale language embedding network for REC. Our method adopts the proposal-free structure, which directly feeds fused visual-language features into a detection head to predict the bounding box of the target. In the fusion process, we propose a grid fusion module and a grid-context fusion module to compute the similarity between language features and visual features in different size regions. Meanwhile, we extra add fully interacted vision-language information and position information to strength the feature fusion. This novel fusion strategy can help to utilize context flexibly therefore the network can deal with varied expressions, especially expressions that describe objects by things around. Our proposed method outperforms the state-of-the-art methods on Refcoco, Refcoco+ and Refcocog datasets. Taijin Zhao, Hongliang Li 0001, Heqian Qiu, Qingbo Wu 0001, King Ngi Ngan |
MMAsia | 1 |