VLDB 2026 Research / reviewers in the wild / expert
Xuelin Zhu
dblp:205/7394
· DBLP profile ↗
19ranked-venue papers
8as first author
16since 2021 · last 2026
0000-0001-7676-2843ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 7 first-author · 14 since 2021Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021Computer networks · 3 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AutoIT: Automated Image Tagging with Random Perturbation
Xuelin Zhu, Jianshu Li, Dongqi Tang, Jiawei Ge 0002, Bo Liu 0004, Jiuxin Cao |
Int. J. Comput. Vis. | 1 |
| 2026 | Multi-Label Image Classification via Contrastive Co-Occurrence LearningabstractMulti-label image classification is an essential task in computer vision that aims to identify multiple objects in images. Recently, there has been growing research interest in modeling the relationships between labels to enhance label representation learning. An intuitive approach is to train a network to estimate label co-occurrence probabilities in a supervised manner, which are then leveraged to guide the interactions between label representations. However, the extreme sparsity of label co-occurrence signals poses substantial challenges. To address this issue, we commence by examining the potential interaction behaviors between label representations under the guidance of ground-truth label co-occurrence signals. Inspired by our findings, a novel contrastive learning mechanism is crafted to mimic and enhance such behaviors, facilitating effective label representation interactions without relying on explicit supervision from label co-occurrence signals. Based on this, we develop a pioneering contrastive co-occurrence learning framework, which operates on the instance-level label co-occurrence graph for multi-label image classification. This framework involves sequential processes of label representation learning followed by co-occurrence perception learning. Cross-entropy loss for label classification learning and contrastive loss for co-occurrence perception learning are used to jointly optimize the entire framework end-to-end. In this way, label representations can interact effectively, fully perceiving their co-occurrence relationships at the instance level, thereby significantly improving the performance in label recognition. Extensive experiments on public benchmarks demonstrate the superiority of the proposed framework in multi-label image classification. Codes are available on https://github.com/jasonseu/CoCo. Xuelin Zhu, Dongqi Tang, Jiawei Ge 0002, Bo Liu 0004, Jiuxin Cao |
IEEE Trans. Image Process. | 1 |
| 2025 | MambaMl: Exploring State Space Models for Multi-Label Image Classification
Xuelin Zhu, Jiuxin Cao, Bing Wang 0013 |
ICCV | 1 |
| 2025 | Denoise-then-Retrieve: Text-Conditioned Video Denoising for Video Moment RetrievalabstractCurrent text-driven Video Moment Retrieval (VMR) methods encode all video clips, including irrelevant ones, disrupting multimodal alignment and hindering optimization. To this end, we propose a denoise-then-retrieve paradigm that explicitly filters text-irrelevant clips from videos and then retrieves the target moment using purified multimodal representations. Following this paradigm, we introduce the Denoise-then-Retrieve Network (DRNet), comprising Text-Conditioned Denoising (TCD) and Text-Reconstruction Feedback (TRF) modules. TCD integrates cross-attention and structured state space blocks to dynamically identify noisy clips and produce a noise mask to purify multimodal video representations. TRF further distills a single query embedding from purified video representations and aligns it with the text embedding, serving as auxiliary supervision for denoising during training. Finally, we perform conditional retrieval using text embeddings on purified video representations for accurate VMR. Experiments on Charades-STA and QVHighlights demonstrate that our approach surpasses state-of-the-art methods on all metrics. Furthermore, our denoise-then-retrieve paradigm is adaptable and can be seamlessly integrated into advanced VMR models to boost performance. Jiuxin Cao, Bo Miao, Zhiheng Fu, Xuelin Zhu, Jiawei Ge 0002, Bo Liu 0004, Mehwish Nasim, Ajmal Mian |
IJCAI | 5 |
| 2025 | MSCI: Addressing CLIP's Inherent Limitations for Compositional Zero-Shot LearningabstractCompositional Zero-Shot Learning (CZSL) aims to recognize unseen state-object combinations by leveraging known combinations. Existing studies basically rely on the cross-modal alignment capabilities of CLIP but tend to overlook its limitations in capturing fine-grained local features, which arise from its architectural and training paradigm. To address this issue, we propose a Multi-Stage Cross-modal Interaction (MSCI) model that effectively explores and utilizes intermediate-layer information from CLIP's visual encoder. Specifically, we design two self-adaptive aggregators to extract local information from low-level visual features and integrate global information from high-level visual features, respectively. These key information are progressively incorporated into textual representations through a stage-by-stage interaction mechanism, significantly enhancing the model’s perception capability for fine-grained local visual information. Additionally, MSCI dynamically adjusts the attention weights between global and local visual information based on different combinations, as well as different elements within the same combination, allowing it to flexibly adapt to diverse scenarios. Experiments on three widely used datasets fully validate the effectiveness and superiority of the proposed model. Data and code are available at https://github.com/ltpwy/MSCI. Xuelin Zhu, Yicong Li 0016 |
IJCAI | 3 |
| 2025 | Gen4Track: A Tuning-free Data Augmentation Framework via Self-correcting Diffusion Model for Vision-Language TrackingabstractThe performance of current Vision-Language Tracking (VLT) models is constrained by the limited diversity and quantity of labeled data. Compared to constructing large-scale datasets, data augmentation offers a more cost-saving strategy for VLT by synthesizing new samples from existing data, rather than generating them from scratch. However, conventional techniques like rotation and flipping may disrupt scene composition, causing conflicts between visual layouts and textual annotations. Recent advances in generative models have inspired the use of synthetic videos for data augmentation. Yet, existing approaches fail to address the core concerns of data augmentation in VLT (shown in Fig. 1)-target location accuracy, text-video consistency, and video content coherency. To bridge the gap, we propose Gen4Track, a tuning-free data augmentation framework that leverages the self-correcting mechanism to dynamically generate high-quality video data with annotations. Our approach involves (1) optimizing the attention calculations in a frozen text-to-image diffusion model to synthesize coherent videos that satisfy specific conditions (e.g., spatial location, category, color, and style), and (2) implementing a self-correcting mechanism based on a Large Language Model (LLM) to improve text-video consistency. During video augmentation, we propose content-coherent self-attention and location-enhanced cross-attention mechanisms, ensuring that image-level editings are accurately and coherently propagated throughout the video. Then, with the goal of maximizing text-video consistency, we iteratively refine the augmentation instruction with our designed self-correcting mechanism for a more aligned video. Extensive experiments validate that Gen4Track significantly boosts the performance of SOTA VLT models (achieving improvements of up to 3.2% in SUC and 3.5% in PRE), opening a new chapter of training Vision-Language trackers with synthetic videos rather than manually annotated data. Jiawei Ge 0002, Xin-Yu Zhang 0027, Jiuxin Cao, Xuelin Zhu, Qingqing Gao, Biwei Cao, Kun Wang 0057, Chang Liu 0113, Bo Liu 0004, Chen Feng 0028, Ioannis Patras |
ACM Multimedia | 4 |
| 2025 | Context-Enhanced Video Moment Retrieval With Large Language ModelsabstractCurrent methods for Video Moment Retrieval (VMR) struggle to align complex situations involving specific environmental details, character descriptions, and action narratives. To tackle this issue, we propose a Large Language Model-guided Moment Retrieval (LMR) approach that employs the extensive knowledge of Large Language Models (LLMs) to improve video context representation as well as cross-modal alignment, facilitating accurate localization of target moments. Specifically, LMR introduces a context enhancement technique with LLMs to generate crucial target-related context semantics. These semantics are integrated with visual features for producing discriminative video representations. Finally, a language-conditioned transformer is designed to decode free-form language queries, on the fly, using aligned video representations for moment retrieval. Extensive experiments demonstrate that LMR achieves state-of-the-art results, outperforming the nearest competitor by up to 3.28% and 4.06% on the challenging QVHighlights and Charades-STA benchmarks, respectively. More importantly, the performance gains are significantly higher for localization of complex queries. Bo Miao, Jiuxin Cao, Xuelin Zhu, Jiawei Ge 0002, Bo Liu 0004, Mehwish Nasim, Ajmal Mian |
IEEE Trans. Multim. | 4 |
| 2025 | Beyond Visual Cues: Synchronously Exploring Target-Centric Semantics for Vision-Language TrackingabstractSingle object tracking aims to locate one specific target in video sequences, given its initial state. Classical trackers rely solely on visual cues, restricting their ability to handle challenges such as appearance variations, ambiguity, and distractions. Hence, Vision-Language Tracking (VLT) has emerged as a promising approach, incorporating language descriptions to directly provide high-level semantics and enhance tracking performance. However, current Vision-Language (VL) trackers have not fully exploited the power of multi-modal learning, as they suffer from limitations such as heavily relying on off-the-shelf backbones for feature extraction, ineffective asynchronous fusion designs, and the absence of VL-related loss functions for optimizing multi-modal representation. Consequently, we present a novel tracker that progressively explores target-centric semantics for VLT. Specifically, we propose the first Synchronous Learning Backbone (SLB) for VLT, which consists of two novel modules: the Target Enhance Module (TEM) and the Semantic-Aware Module (SAM). These modules together ensure the multi-modal feature extraction and interaction at the same pace, facilitating the tracker to synchronously perceive target-related semantics from both visual and textual modalities. Moreover, we devise the dense matching loss to further strengthen multi-modal representation learning. Extensive experiments on VLT datasets demonstrate the superiority and effectiveness of our methods. Jiawei Ge 0002, Jiuxin Cao, Xiangmei Chen, Xuelin Zhu, Chang Liu 0113, Kun Wang 0057, Bo Liu 0004 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2025 | Query-Based Knowledge Sharing for Open-Vocabulary Multi-Label ClassificationabstractIdentifying labels that are unseen during training, known as multi-label zero-shot learning, is a non-trivial task in computer vision. Recent studies have increasingly focused on utilizing vision-language pre-training (VLP) models to recognize unseen labels in an open-vocabulary manner. However, these approaches like knowledge distillation have offered only modest performance gains. The challenge of fully harnessing the potential of VLP models for effective multi-label zero-shot learning remains open. In this work, an advanced query-based knowledge sharing framework is proposed to explore the multi-modal knowledge from VLP models for open-vocabulary multi-label classification. Specifically, we introduce a set of label-agnostic query tokens that are designed to capture essential and informative visual knowledge from input images. These tokens are subsequently shared across all labels, allowing them to select pertinent one as visual clues for accurate recognition. Then, by integrating the pre-trained knowledge of VLP models, these query tokens, trained on seen labels, can be efficiently generalized to the recognition of unseen labels. Additionally, we reformulate ranking learning into a form of classification to enable the magnitude of feature vectors for prediction, which significantly benefits label recognition. Experiment results show that our framework outperforms state-of-the-art methods in multi-label zero-shot learning task by a significant margin, reaching 4.2% and 2.4% in mAP on the NUS-WIDE and Open Images datasets, respectively. Code and models are available at https://github.com/jasonseu/QKS . Xuelin Zhu, Dongqi Tang, Jiawei Ge 0002, Bo Liu 0004, Jiuxin Cao |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2024 | Consistencies are All You Need for Semi-supervised Vision-Language TrackingabstractVision-Language Tracking (VLT) requires locating a specific target in video sequences, given a natural language prompt and an initial object box. Despite recent advancements, existing approaches heavily rely on expensive and time-consuming human annotations. To mitigate this limitation, directly generating pseudo labels from raw videos seems to be a straightforward solution; however, it inevitably introduces undesirable noise during the training process. Moreover, we insist that an efficient tracker should excel in tracking the target, regardless of the temporal direction. Building upon these insights, we propose the pioneering semi-supervised learning scheme for VLT task, representing a crucial step towards reducing the dependency on high-quality yet costly labeled data. Specifically, drawing inspiration from the natural attributes of a video (i.e., space, time, and semantics), our approach progressively leverages inherent consistencies from these aspects: (1) Spatially, each frame and any object cropped from it naturally form an image-bbox (bounding box) pair for self-training; (2) Temporally, bidirectional tracking trajectories should exhibit minimal differences; (3) Semantically, the correlation between visual and textual features is expected to remain consistent. Furthermore, the framework is validated with a simple yet effective tracker we devised, named ATTracker (Asymmetrical Transformer Tracker). It modifies the self-attention operation in an asymmetrical way, striving to enhance target-related features while suppressing noise. Extensive experiments confirm that our ATTracker serves as a robust baseline, outperforming fully supervised base trackers. By unveiling the potential of learning with limited annotations, this study aims to attract attention and pave the way for Semi-supervised Vision-Language Tracking (SS-VLT). Jiawei Ge 0002, Jiuxin Cao, Xuelin Zhu, Xin-Yu Zhang 0027, Chang Liu 0113, Kun Wang 0057, Bo Liu 0004 |
ACM Multimedia | 3 |
| 2024 | Enhancing Micro-Video Venue Recognition via Multi-Modal and Multi-Granularity Object RelationsabstractMicro-video venue recognition aims to predict the venue category where a micro-video was filmed. Different from traditional long videos which contain rich temporal context, venue prediction for micro-videos is difficult due to its limited duration (generally within 6s). The existing works usually extract features of each modality from a global perspective for prediction, neglecting the semantics carried by local objects. To this end, we propose Multi-Modal and Multi-Granularity Object Relations (M2ORE) to address the above issues, which learns multi-granularity interactive semantics between venues and multimodal semantic objects to help understand venues. Specifically, M2ORE comprises of two modules: it first extract semantic objects of different modalities, i.e. visual objects in keyframes and keywords in texts, and models the affiliation relationship between semantic objects and venues and the co-occurrence relationship among semantic objects, forming a heterogeneous venue-object relation graph. Then, to achieve the interactive semantics between venues and objects from the relation graph, a novel Parallel-Graph Inference Model (Parallel-GIM) is proposed, which updates the representation of nodes through graph propagation and fuse multi-level features (local-global-multimodal) through the devised hierarchical attention mechanism. Finally, the probability distribution of venues can be obtained through a multi-layer perceptron with the comprehensive features of the venue nodes. Extensive experiments on real-world micro-video dataset demonstrate the superiority of the proposed M2ORE. Jiuxin Cao, Xuelin Zhu, Bo Liu 0004 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Semantic-Guided Representation Enhancement for Multi-Label Image ClassificationabstractMulti-label image classification is an essential yet challenging task that requires to recognize multiple objects of images. To this end, recent studies have sought to acquire visual representations for each label by attention models, and then train binary classifiers for prediction. However, these methods have two major drawbacks: 1) They rely heavily on the precise alignments between two modalities, which is still challenging for current attention models; 2) They ignore patch-level representations rich in local object features, which are also of great importance for label recognition. In this paper, we propose a semantic-guided representation enhancement framework, which augments patch-level representations with object-level representations for robust label recognition. Concretely, the proposed framework consists of two significant components: 1) an inter-modal attention module that accounts for coarsely locating object regions and producing object-level representations for each label; 2) an intra-modal attention module that aggregates object representations to enhance patch representations based on their correlations. In this way, both local clues and global glances of objects are fully exploited simultaneously, rather than relying solely on object-level representations obtained by the inter-modal attention, thus improving the performance of label recognition. Experimental results show that our framework outperforms the state-of-the-art methods by 0.5%, 0.6%, 0.7% and 0.8% in mAP on Pascal VOC 2007, Microsoft COCO, NUS-WIDE and Visual Genome datasets, respectively. Codes and models are available on https://github.com/jasonseu/SGRE. Xuelin Zhu, Jianshu Li, Jiuxin Cao, Dongqi Tang, Bo Liu 0004 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | Scene-Aware Label Graph Learning for Multi-Label Image ClassificationabstractMulti-label image classification refers to assigning a set of labels for an image. One of the main challenges of this task is how to effectively capture the correlation among labels. Existing studies on this issue mostly rely on the statistical label co-occurrence or semantic similarity of labels. However, an important fact is ignored that the co-occurrence of labels is closely related with image scenes (indoor, outdoor, etc.), which is a vital characteristic in multi-label image classification. In this paper, a novel scene-aware label graph learning framework is proposed, which is capable of learning visual representations for labels while fully perceiving their co-occurrence relationships under variable scenes. Specifically, our framework is able to detect scene categories of images without relying on manual annotations, and keeps track of the co-occurring labels by maintaining a global co-occurrence matrix for each scene category throughout the whole training phase. These scene-independent co-occurrence matrices are further employed to guide the interactions among label representations in a graph propagation manner towards accurate label prediction. Extensive experiments on public benchmarks demonstrate the superiority of our framework. Xuelin Zhu, Jiawei Ge 0002, Bo Liu 0004, Jiuxin Cao |
ICCV | 1 |
| 2023 | Exploring Visual Pre-training for Robot Manipulation: Datasets, Models and MethodsabstractVisual pre-training with large-scale real-world data has made great progress in recent years, showing great potential in robot learning with pixel observations. However, the recipes of visual pre-training for robot manipulation tasks are yet to be built. In this paper, we thoroughly investigate the effects of visual pre-training strategies on robot manipulation tasks from three fundamental perspectives: pre-training datasets, model architectures and training methods. Several significant experimental findings are provided that are beneficial for robot learning. Further, we propose a visual pre-training scheme for robot manipulation termed Vi-PRoM, which combines self-supervised learning and supervised learning. Concretely, the former employs contrastive learning to acquire underlying patterns from large-scale unlabeled data, while the latter aims learning visual semantics and temporal dynamics. Extensive experiments on robot manipulations in various simulation environments and the real robot demonstrate the superiority of the proposed scheme. Videos and more details can be found on https://explore-pretrain-robot.github.io. Ya Jing, Xuelin Zhu, Xingbin Liu, Qie Sima, Taozheng Yang, Yunhai Feng, Tao Kong |
IROS | 2 |
| 2023 | Real-time anomaly detection on surveillance video with two-stream spatio-temporal generative model
Jiuxin Cao, Bo Liu 0004, Xuelin Zhu |
Multim. Syst. | 5 |
| 2022 | Two-Stream Transformer for Multi-Label Image ClassificationabstractMulti-label image classification is a fundamental yet challenging task in computer vision that aims to identify multiple objects from a given image. Recent studies on this task mainly focus on learning cross-modal interactions between label semantics and high-level visual representations via an attention operation. However, these one-shot attention based approaches generally perform poorly in establishing accurate and robust alignments between vision and text due to the acknowledged semantic gap. In this paper, we propose a two-stream transformer (TSFormer) learning framework, in which the spatial stream focuses on extracting patch features with a global perception, while the semantic stream aims to learn vision-aware label semantics as well as their correlations via a multi-shot attention mechanism. Specifically, in each layer of TSFormer, a cross-modal attention module is developed to aggregate visual features from spatial stream into semantic stream and update label semantics via a residual connection. In this way, the semantic gap between two streams gradually narrows as the procedure progresses layer by layer, allowing the semantic stream to produce sophisticated visual representations for each label towards accurate label recognition. Extensive experiments on three visual benchmarks, including Pascal VOC 2007, Microsoft COCO and NUS-WIDE, consistently demonstrate that our proposed TSFormer achieves state-of-the-art performance on the multi-label image classification task. Xuelin Zhu, Jiuxin Cao, Jiawei Ge 0002, Bo Liu 0004 |
ACM Multimedia | 1 |
| 2019 | Joint Visual-Textual Sentiment Analysis Based on Cross-Modality Attention Mechanism
Xuelin Zhu, Biwei Cao, Bo Liu 0004, Jiuxin Cao |
MMM (1) | 1 |
| 2018 | Community Discovery Based on Social Relations and Temporal-Spatial Topics in LBSNs
Jiuxin Cao, Xuelin Zhu, Bo Liu 0004 |
PAKDD (3) | 3 |
| 2018 | Effective fine-grained location prediction based on user check-in pattern in LBSNs
Jiuxin Cao, Xuelin Zhu, Renjun Lv, Bo Liu 0004 |
J. Netw. Comput. Appl. | 3 |