Jiawei Ge 0002

dblp:250/3739-2 · also Jia-Wei Ge 0002 · DBLP profile ↗
← Back
17ranked-venue papers
6as first author
17since 2021 · last 2026
0000-0001-7268-7815ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 3 first-author · 11 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Computer networks · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 AutoIT: Automated Image Tagging with Random Perturbation
Xuelin Zhu, Jianshu Li, Dongqi Tang, Jiawei Ge 0002, Bo Liu 0004, Jiuxin Cao
Int. J. Comput. Vis.5
2026 Multi-Label Image Classification via Contrastive Co-Occurrence Learning
abstract
Multi-label image classification is an essential task in computer vision that aims to identify multiple objects in images. Recently, there has been growing research interest in modeling the relationships between labels to enhance label representation learning. An intuitive approach is to train a network to estimate label co-occurrence probabilities in a supervised manner, which are then leveraged to guide the interactions between label representations. However, the extreme sparsity of label co-occurrence signals poses substantial challenges. To address this issue, we commence by examining the potential interaction behaviors between label representations under the guidance of ground-truth label co-occurrence signals. Inspired by our findings, a novel contrastive learning mechanism is crafted to mimic and enhance such behaviors, facilitating effective label representation interactions without relying on explicit supervision from label co-occurrence signals. Based on this, we develop a pioneering contrastive co-occurrence learning framework, which operates on the instance-level label co-occurrence graph for multi-label image classification. This framework involves sequential processes of label representation learning followed by co-occurrence perception learning. Cross-entropy loss for label classification learning and contrastive loss for co-occurrence perception learning are used to jointly optimize the entire framework end-to-end. In this way, label representations can interact effectively, fully perceiving their co-occurrence relationships at the instance level, thereby significantly improving the performance in label recognition. Extensive experiments on public benchmarks demonstrate the superiority of the proposed framework in multi-label image classification. Codes are available on https://github.com/jasonseu/CoCo.
Xuelin Zhu, Dongqi Tang, Jiawei Ge 0002, Bo Liu 0004, Jiuxin Cao
IEEE Trans. Image Process.5
2025 SKL-CLIP: Learning Skeleton-Based Action Representations via Language Supervision
abstract
CLIP, widely used in multimodal learning, excels due to its large-scale image-text pretraining. However, applying CLIP-like architectures to skeleton-based action representation learning presents challenges due to the incompatible non-visual data structure and the limited scale of skeleton datasets, which hinder robust generalization. To address these issues, we propose SKL-CLIP, a framework that incorporates Supervised Self-Contrastive Learning to mitigate overfitting and enhance transferable representation learning, Knowledge Distillation from the textual encoder of pretrained CLIP models to preserve generalization while adapting to skeleton-text scenarios, and Multi-Domain Parallel Training to leverage diverse support datasets, improving cross-dataset and zero-shot recognition. Extensive experiments on NTU and PKU datasets demonstrate that SKL-CLIP significantly advances skeleton-based action representation, achieving state-of-the-art performance across fully-supervised, cross-data unsupervised domain adaptation, and zero-shot tasks.
Kun Wang 0057, Jiuxin Cao, Jiawei Ge 0002, Chang Liu 0113, Bo Liu 0004
ICME3
2025 Denoise-then-Retrieve: Text-Conditioned Video Denoising for Video Moment Retrieval
abstract
Current text-driven Video Moment Retrieval (VMR) methods encode all video clips, including irrelevant ones, disrupting multimodal alignment and hindering optimization. To this end, we propose a denoise-then-retrieve paradigm that explicitly filters text-irrelevant clips from videos and then retrieves the target moment using purified multimodal representations. Following this paradigm, we introduce the Denoise-then-Retrieve Network (DRNet), comprising Text-Conditioned Denoising (TCD) and Text-Reconstruction Feedback (TRF) modules. TCD integrates cross-attention and structured state space blocks to dynamically identify noisy clips and produce a noise mask to purify multimodal video representations. TRF further distills a single query embedding from purified video representations and aligns it with the text embedding, serving as auxiliary supervision for denoising during training. Finally, we perform conditional retrieval using text embeddings on purified video representations for accurate VMR. Experiments on Charades-STA and QVHighlights demonstrate that our approach surpasses state-of-the-art methods on all metrics. Furthermore, our denoise-then-retrieve paradigm is adaptable and can be seamlessly integrated into advanced VMR models to boost performance.
Jiuxin Cao, Bo Miao, Zhiheng Fu, Xuelin Zhu, Jiawei Ge 0002, Bo Liu 0004, Mehwish Nasim, Ajmal Mian
IJCAI6
2025 Gen4Track: A Tuning-free Data Augmentation Framework via Self-correcting Diffusion Model for Vision-Language Tracking
abstract
The performance of current Vision-Language Tracking (VLT) models is constrained by the limited diversity and quantity of labeled data. Compared to constructing large-scale datasets, data augmentation offers a more cost-saving strategy for VLT by synthesizing new samples from existing data, rather than generating them from scratch. However, conventional techniques like rotation and flipping may disrupt scene composition, causing conflicts between visual layouts and textual annotations. Recent advances in generative models have inspired the use of synthetic videos for data augmentation. Yet, existing approaches fail to address the core concerns of data augmentation in VLT (shown in Fig. 1)-target location accuracy, text-video consistency, and video content coherency. To bridge the gap, we propose Gen4Track, a tuning-free data augmentation framework that leverages the self-correcting mechanism to dynamically generate high-quality video data with annotations. Our approach involves (1) optimizing the attention calculations in a frozen text-to-image diffusion model to synthesize coherent videos that satisfy specific conditions (e.g., spatial location, category, color, and style), and (2) implementing a self-correcting mechanism based on a Large Language Model (LLM) to improve text-video consistency. During video augmentation, we propose content-coherent self-attention and location-enhanced cross-attention mechanisms, ensuring that image-level editings are accurately and coherently propagated throughout the video. Then, with the goal of maximizing text-video consistency, we iteratively refine the augmentation instruction with our designed self-correcting mechanism for a more aligned video. Extensive experiments validate that Gen4Track significantly boosts the performance of SOTA VLT models (achieving improvements of up to 3.2% in SUC and 3.5% in PRE), opening a new chapter of training Vision-Language trackers with synthetic videos rather than manually annotated data.
Jiawei Ge 0002, Xin-Yu Zhang 0027, Jiuxin Cao, Xuelin Zhu, Qingqing Gao, Biwei Cao, Kun Wang 0057, Chang Liu 0113, Bo Liu 0004, Chen Feng 0028, Ioannis Patras
ACM Multimedia1
2025 FSD-GAN: Generative Adversarial Training for Face Swap Detection via the Latent Noise Fingerprint
Jiawei Ge 0002, Jiuxin Cao, Zhixiang Zhao, Bo Liu 0004
J. Comput. Sci. Technol.1
2025 Causality-inspired representation learning for weakly supervised skeleton-based action recognition
Kun Wang 0057, Jiuxin Cao, Jiawei Ge 0002, Chang Liu 0113, Bo Liu 0004
Knowl. Based Syst.3
2025 Context-Enhanced Video Moment Retrieval With Large Language Models
abstract
Current methods for Video Moment Retrieval (VMR) struggle to align complex situations involving specific environmental details, character descriptions, and action narratives. To tackle this issue, we propose a Large Language Model-guided Moment Retrieval (LMR) approach that employs the extensive knowledge of Large Language Models (LLMs) to improve video context representation as well as cross-modal alignment, facilitating accurate localization of target moments. Specifically, LMR introduces a context enhancement technique with LLMs to generate crucial target-related context semantics. These semantics are integrated with visual features for producing discriminative video representations. Finally, a language-conditioned transformer is designed to decode free-form language queries, on the fly, using aligned video representations for moment retrieval. Extensive experiments demonstrate that LMR achieves state-of-the-art results, outperforming the nearest competitor by up to 3.28% and 4.06% on the challenging QVHighlights and Charades-STA benchmarks, respectively. More importantly, the performance gains are significantly higher for localization of complex queries.
Bo Miao, Jiuxin Cao, Xuelin Zhu, Jiawei Ge 0002, Bo Liu 0004, Mehwish Nasim, Ajmal Mian
IEEE Trans. Multim.5
2025 Beyond Visual Cues: Synchronously Exploring Target-Centric Semantics for Vision-Language Tracking
abstract
Single object tracking aims to locate one specific target in video sequences, given its initial state. Classical trackers rely solely on visual cues, restricting their ability to handle challenges such as appearance variations, ambiguity, and distractions. Hence, Vision-Language Tracking (VLT) has emerged as a promising approach, incorporating language descriptions to directly provide high-level semantics and enhance tracking performance. However, current Vision-Language (VL) trackers have not fully exploited the power of multi-modal learning, as they suffer from limitations such as heavily relying on off-the-shelf backbones for feature extraction, ineffective asynchronous fusion designs, and the absence of VL-related loss functions for optimizing multi-modal representation. Consequently, we present a novel tracker that progressively explores target-centric semantics for VLT. Specifically, we propose the first Synchronous Learning Backbone (SLB) for VLT, which consists of two novel modules: the Target Enhance Module (TEM) and the Semantic-Aware Module (SAM). These modules together ensure the multi-modal feature extraction and interaction at the same pace, facilitating the tracker to synchronously perceive target-related semantics from both visual and textual modalities. Moreover, we devise the dense matching loss to further strengthen multi-modal representation learning. Extensive experiments on VLT datasets demonstrate the superiority and effectiveness of our methods.
Jiawei Ge 0002, Jiuxin Cao, Xiangmei Chen, Xuelin Zhu, Chang Liu 0113, Kun Wang 0057, Bo Liu 0004
ACM Trans. Multim. Comput. Commun. Appl.1
2025 Dual-Domain Triple Contrast for Cross-Dataset Skeleton-Based Action Recognition
abstract
Skeleton-based Action Recognition (SAR) is widely recognized for its robustness and efficiency in human action analysis, but its performance in cross-dataset tasks has been limited due to domain shifts between different datasets. To address this challenge, current methods typically approach cross-dataset SAR as an Unsupervised Domain Adaptation (UDA) task, which is tackled using domain adaptation or self-supervised learning strategies. In this article, we propose a Dual-Domain Triple Contrast (D2TC) framework for cross-dataset SAR under the UDA setting. Unlike existing UDA methods that either focus on a single strategy or superficially combine strategies, our D2TC leverages contrastive learning to integrate both strategies into a unified framework. It performs three types of contrastive learning: Self-Supervised Contrastive Learning, Supervised Contrastive Learning, and UDA with Contrastive Learning, across both source and target domains. The triple contrasts go beyond mere summation, effectively bridging the domain gap and enhancing the model’s representational capacity. Additionally, we introduce multi-modal ensemble contrast and extreme skeleton augmentation methods to further enhance the skeleton-based representation learning. Extensive experiments on six cross-dataset settings validate the superiority of our D2TC framework over state-of-the-art methods, demonstrating its effectiveness in reducing domain discrepancies and improving cross-dataset SAR performance. The codes are available on https://github.com/KennCoder7/DualDomainTripleContrast .
Kun Wang 0057, Jiuxin Cao, Jiawei Ge 0002, Chang Liu 0113, Bo Liu 0004
ACM Trans. Multim. Comput. Commun. Appl.3
2025 Query-Based Knowledge Sharing for Open-Vocabulary Multi-Label Classification
abstract
Identifying labels that are unseen during training, known as multi-label zero-shot learning, is a non-trivial task in computer vision. Recent studies have increasingly focused on utilizing vision-language pre-training (VLP) models to recognize unseen labels in an open-vocabulary manner. However, these approaches like knowledge distillation have offered only modest performance gains. The challenge of fully harnessing the potential of VLP models for effective multi-label zero-shot learning remains open. In this work, an advanced query-based knowledge sharing framework is proposed to explore the multi-modal knowledge from VLP models for open-vocabulary multi-label classification. Specifically, we introduce a set of label-agnostic query tokens that are designed to capture essential and informative visual knowledge from input images. These tokens are subsequently shared across all labels, allowing them to select pertinent one as visual clues for accurate recognition. Then, by integrating the pre-trained knowledge of VLP models, these query tokens, trained on seen labels, can be efficiently generalized to the recognition of unseen labels. Additionally, we reformulate ranking learning into a form of classification to enable the magnitude of feature vectors for prediction, which significantly benefits label recognition. Experiment results show that our framework outperforms state-of-the-art methods in multi-label zero-shot learning task by a significant margin, reaching 4.2% and 2.4% in mAP on the NUS-WIDE and Open Images datasets, respectively. Code and models are available at https://github.com/jasonseu/QKS .
Xuelin Zhu, Dongqi Tang, Jiawei Ge 0002, Bo Liu 0004, Jiuxin Cao
ACM Trans. Multim. Comput. Commun. Appl.4
2024 Consistencies are All You Need for Semi-supervised Vision-Language Tracking
abstract
Vision-Language Tracking (VLT) requires locating a specific target in video sequences, given a natural language prompt and an initial object box. Despite recent advancements, existing approaches heavily rely on expensive and time-consuming human annotations. To mitigate this limitation, directly generating pseudo labels from raw videos seems to be a straightforward solution; however, it inevitably introduces undesirable noise during the training process. Moreover, we insist that an efficient tracker should excel in tracking the target, regardless of the temporal direction. Building upon these insights, we propose the pioneering semi-supervised learning scheme for VLT task, representing a crucial step towards reducing the dependency on high-quality yet costly labeled data. Specifically, drawing inspiration from the natural attributes of a video (i.e., space, time, and semantics), our approach progressively leverages inherent consistencies from these aspects: (1) Spatially, each frame and any object cropped from it naturally form an image-bbox (bounding box) pair for self-training; (2) Temporally, bidirectional tracking trajectories should exhibit minimal differences; (3) Semantically, the correlation between visual and textual features is expected to remain consistent. Furthermore, the framework is validated with a simple yet effective tracker we devised, named ATTracker (Asymmetrical Transformer Tracker). It modifies the self-attention operation in an asymmetrical way, striving to enhance target-related features while suppressing noise. Extensive experiments confirm that our ATTracker serves as a robust baseline, outperforming fully supervised base trackers. By unveiling the potential of learning with limited annotations, this study aims to attract attention and pave the way for Semi-supervised Vision-Language Tracking (SS-VLT).
Jiawei Ge 0002, Jiuxin Cao, Xuelin Zhu, Xin-Yu Zhang 0027, Chang Liu 0113, Kun Wang 0057, Bo Liu 0004
ACM Multimedia1
2024 GAL: combining global and local contexts for interpersonal relation extraction toward document-level Chinese text
Jiawei Ge 0002, Jiuxin Cao, Yingxing Bao, Biwei Cao, Bo Liu 0004
Neural Comput. Appl.1
2024 Correction: GAL: combining global and local contexts for interpersonal relation extraction toward document-level Chinese text
Jiawei Ge 0002, Jiuxin Cao, Yingxing Bao, Biwei Cao, Bo Liu 0004
Neural Comput. Appl.1
2024 A Multi-Agent System for Fine-Grained Opinion Dynamics Analysis in Online Social Networks
abstract
The influence and dissemination of users’ opinions are the essence of the opinion dynamics in online social networks (OSNs). Understanding the process of users’ opinion formation and propagation can provide better service for public opinion monitoring and product advertising. However, most opinion dynamics research typically do not distinguish between user opinion formation and propagation, instead focusing on the process of mass opinion polarization. This article presents a multiagent system (MAS) to analyze fine-grained opinion dynamics (FOD) by agent-based modeling (ABM). At the macro level, the MAS-FOD system we designed can produce a statistical analysis of public opinion evolution based on different parameter constellations and analyze the changes in personalized agent opinions. At the micro level, the agent-based model is presented to describe the single user in OSNs as an individualized agent in MAS-FOD. Specifically, we propose two mechanisms, agent-based opinion formation (AOF) mechanism and agent-based opinion propagation (AOP) mechanism, for agents to form and disseminate opinions, respectively. Meanwhile, multidimensional social influence features mining from self, local neighbors, and global topic community are defined and used in two mechanisms to train personalized agents to self-adapt to the MAS-FOD system. We demonstrate the rationality and effectiveness of the MAS-FOD system from two types of experiments: empirical analysis and simulation analysis. The former is driven by empirical social network data to analyze the predictive performance of AOF; the latter simulates the public opinion propagation through the AOP mechanism in the MAS-FOD system. The results demonstrate that: 1) the AOF outperforms SOTA methods in the accuracy of agent opinion prediction; 2) the dissemination effect based on the AOP is more closer to the actual evolution trend of public opinion than the baseline; and 3) the MAS-FOD system can perform fine-grained analysis of opinion dynamics by adjusting different parameters.
Huiyu Min, Jiuxin Cao, Jiawei Ge 0002, Bo Liu 0004
IEEE Trans. Comput. Soc. Syst.3
2023 Scene-Aware Label Graph Learning for Multi-Label Image Classification
abstract
Multi-label image classification refers to assigning a set of labels for an image. One of the main challenges of this task is how to effectively capture the correlation among labels. Existing studies on this issue mostly rely on the statistical label co-occurrence or semantic similarity of labels. However, an important fact is ignored that the co-occurrence of labels is closely related with image scenes (indoor, outdoor, etc.), which is a vital characteristic in multi-label image classification. In this paper, a novel scene-aware label graph learning framework is proposed, which is capable of learning visual representations for labels while fully perceiving their co-occurrence relationships under variable scenes. Specifically, our framework is able to detect scene categories of images without relying on manual annotations, and keeps track of the co-occurring labels by maintaining a global co-occurrence matrix for each scene category throughout the whole training phase. These scene-independent co-occurrence matrices are further employed to guide the interactions among label representations in a graph propagation manner towards accurate label prediction. Extensive experiments on public benchmarks demonstrate the superiority of our framework.
Xuelin Zhu, Jiawei Ge 0002, Bo Liu 0004, Jiuxin Cao
ICCV4
2022 Two-Stream Transformer for Multi-Label Image Classification
abstract
Multi-label image classification is a fundamental yet challenging task in computer vision that aims to identify multiple objects from a given image. Recent studies on this task mainly focus on learning cross-modal interactions between label semantics and high-level visual representations via an attention operation. However, these one-shot attention based approaches generally perform poorly in establishing accurate and robust alignments between vision and text due to the acknowledged semantic gap. In this paper, we propose a two-stream transformer (TSFormer) learning framework, in which the spatial stream focuses on extracting patch features with a global perception, while the semantic stream aims to learn vision-aware label semantics as well as their correlations via a multi-shot attention mechanism. Specifically, in each layer of TSFormer, a cross-modal attention module is developed to aggregate visual features from spatial stream into semantic stream and update label semantics via a residual connection. In this way, the semantic gap between two streams gradually narrows as the procedure progresses layer by layer, allowing the semantic stream to produce sophisticated visual representations for each label towards accurate label recognition. Extensive experiments on three visual benchmarks, including Pascal VOC 2007, Microsoft COCO and NUS-WIDE, consistently demonstrate that our proposed TSFormer achieves state-of-the-art performance on the multi-label image classification task.
Xuelin Zhu, Jiuxin Cao, Jiawei Ge 0002, Bo Liu 0004
ACM Multimedia3