Zuo Cao

dblp:123/1844 · DBLP profile ↗
← Back
14ranked-venue papers
0as first author
14since 2021 · last 2025
0000-0001-9446-5105ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 9 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021
YearPublicationVenuePosition
2025 Intrinsic Consistency Preservation With Adaptively Reliable Samples for Source-Free Domain Adaptation
abstract
Unsupervised domain adaptation (UDA) aims to alleviate the domain shift by transferring knowledge learned from a labeled source dataset to an unlabeled target domain. Although UDA has seen promising progress recently, it requires access to data from both domains, making it problematic in source data-absent scenarios. In this article, we investigate a practical task source-free domain adaptation (SFDA) that alleviates the limitations of the widely studied UDA in simultaneously acquiring source and target data. In addition, we further study the imbalanced SFDA (ISFDA) problem, which addresses the intra-domain class imbalance and inter-domain label shift in SFDA. We observe two key issues in SFDA that: 1) target data form clusters in the representation space regardless of whether the target data points are aligned with the source classifier and 2) target samples with higher classification confidence are more reliable and have less variation in their classification confidence during adaptation. Motivated by these observations, we propose a unified method, named intrinsic consistency preservation with adaptively reliable samples (ICPR), to jointly cope with SFDA and ISFDA. Specifically, ICPR first encourages the intrinsic consistency in the predictions of neighbors for unlabeled samples with weak augmentation (standard flip-and-shift), regardless of their reliability. ICPR then generates strongly augmented views specifically for adaptively selected reliable samples and is trained to fix the intrinsic consistency between weakly and strongly augmented views of the same image concerning predictions of neighbors and their own. Additionally, we propose to use a prototype-like classifier to avoid the classification confusion caused by severe intra-domain class imbalance and inter-domain label shift. We demonstrate the effectiveness and general applicability of ICPR on six benchmarks of both SFDA and ISFDA tasks. The reproducible code of our proposed ICPR method is available at https://github.com/CFM-MSG/Code_ICPR.
Abdulmotaleb El Saddik, Xing Xu 0001, Dongshuai Li, Zuo Cao, Heng Tao Shen
IEEE Trans. Neural Networks Learn. Syst.5
2024 SDN: Semantic Decoupling Network for Temporal Language Grounding
abstract
Temporal language grounding (TLG) is one of the most challenging cross-modal video understanding tasks, which aims at retrieving the most relevant video segment from an untrimmed video according to a natural language sentence. The existing methods can be separated into two dominant types: 1) proposal-based and 2) proposal-free methods, where the former conduct contextual interactions and the latter localizes timestamps flexibly. However, the constant-scale candidates in proposal-based methods limit the localization precision and bring extra computational costs. In contrast, the proposal-free methods perform well on high-precision metrics-based on the fine-grained features but suffer from a lack of coarse-grained interactions, which cause degeneration when the video becomes complex. In this article, we propose a novel framework termed semantic decoupling network (SDN) that combines the advantages of proposal-based and proposal-free methods and overcomes their defects. It contains three key components: 1) semantic decoupling module (SDM); 2) context modeling block (CMB); and 3) semantic cross-level aggregation module (SCAM). By capturing the video-text contexts in multilevel semantics, the SDM and CMB effectively utilize the benefits of proposal-based methods. Meanwhile, the SCAM maintains the merit of proposal-free methods in that it localizes timestamps precisely. The experiments on three challenge datasets, i.e., Charades-STA, TACoS, and ActivityNet-Caption, show that our proposed SDN method significantly outperforms recent state-of-the-art methods, especially the proposal-free methods. Extensive analyses, as well as the implementation code of the proposed SDN method, are provided at https://github.com/CFM-MSG/Code_SDN.
Xun Jiang 0001, Xing Xu 0001, Jingran Zhang, Fumin Shen, Zuo Cao, Heng Tao Shen
IEEE Trans. Neural Networks Learn. Syst.5
2023 Progressive Event Alignment Network for Partial Relevant Video Retrieval
abstract
Currently, most existing text-based video retrieval methods are only adapted to trimmed videos. However, more complicated untrimmed videos are common in multimedia applications nowadays. In this paper, we focus on the Partially Relevant Video Retrieval (PRVR) task that retrieves untrimmed long videos with partial text descriptions. To tackle this challenging problem, we propose a novel method termed Progressive Event Alignment Network (PEAN) to align text queries with local video content progressively. Specifically, it consists of three key components: (1) A Multimodal Representation Module (MRM) that extracts text representations and hierarchical video representations. (2) An Event Searching Module (ESM) that localizes the described video content roughly. (3) An Event Aligning Module (EAM) that aligns text queries and local video content at a fine-grained level. Additionally, we also design a Gaussian-based pooling strategy in both the ESM and EAM, which thoroughly mines the semantic information in representative video frames. The extensive experiments on three PRVR benchmarks demonstrate our proposed PEAN method significantly outperforms current state-of-the-art methods.
Xun Jiang 0001, Xing Xu 0001, Fumin Shen, Zuo Cao
ICME5
2023 Information Selection-based Domain Adaptation from Black-box Predictors
abstract
Unsupervised domain adaptation aims to address the problem of under labeling by performing knowledge transfer between labeled source domains and unlabeled target domains. Despite impressive progress, learning methods that rely on raw data and raw source model parameters all have the potential for privacy leakage in real life. In recent studies, the source model is set up as a black-box model with only inputs and outputs available, and knowledge distillation is introduced to fit the target model. However, the results of knowledge distillation are affected by confusion-prone instances and incorrect predictions of teacher networks, so we propose an Information Selection-based Knowledge Distillation (ISKD) strategy to perform more efficient distillation. We first perform semantic-level optimization of the source model output information through the association of categories and then filter the instance-level information with constructed confidence scores. In addition to this, the introduction of the self-distillation mechanism further improves the model performance. We conduct experiments on three benchmark datasets and obtain state-of-the-art performance.
Xing Xu 0001, Zuo Cao, Gong Zhang 0007
ICME4
2023 Zero-shot Sketch-based Image Retrieval with Adaptive Balanced Discriminability and Generalizability
abstract
Zero-shot sketch-based image retrieval (ZS-SBIR) is a task that learns semantic knowledge and embedding extraction to retrieve similar images using a sketch without any training examples of unseen classes. Existing methods have attempted to address the modal and semantic gaps in ZS-SBIR by using various strategies such as leveraging category linguistic information for improved discriminability and utilizing knowledge distillation to increase the model’s generalizability towards unseen classes. However, these methods fail to consider the importance of discriminability and generalizability in a unified manner. To address this, we propose a novel method called Adaptive Balanced Discriminability and Generalizability (ABDG) for ZS-SBIR. Specifically, our ABDG method utilizes an advanced two-stage knowledge distillation scheme to balance the learning of discriminability and generalizability for each instance. In addition to task-agnostic teacher models to preserve structural information used in existing work, we introduce a task-specific teacher model pre-trained with a classification objective function to emphasize the discriminability property during knowledge distillation. We also employ a novel entropy-based weighting strategy to balance the effects of structural information preservation and classification losses specific to the classification progress of each instance. Furthermore, we use fine-grained semantic relevance to refine the ego predictions of the student model, with the aim of improving its performance as the training objective continues to converge. Experimental results on three benchmark datasets of ZS-SBIR demonstrate that our ABDG method establishes a state-of-the-art performance by balancing the learning of discriminative and generalizable properties.
Xing Xu 0001, Zuo Cao, Gong Zhang 0007, Fumin Shen, Yang Yang 0002
ICMR3
2023 Joint Searching and Grounding: Multi-Granularity Video Content Retrieval
abstract
Text-based video retrieval is a well-studied task aimed at retrieving relevant videos from a large collection in response to a given text query. Most existing TVR works assume that videos are already trimmed and fully relevant to the query thus ignoring that most videos in real-world scenarios are untrimmed and contain massive irrelevant video content. Moreover, as users' queries are only relevant to video events rather than complete videos, it is also more practical to provide specific video events rather than an untrimmed video list. In this paper, we introduce a challenging but more realistic task called Multi-Granularity Video Content Retrieval (MGVCR), which involves retrieving both video files and specific video content with their temporal locations. This task presents significant challenges since it requires identifying and ranking the partial relevance between long videos and text queries under the lack of temporal alignment supervision between the query and relevant moments. To this end, we propose a novel unified framework, termed, Joint Searching and Grounding (JSG). It consists of two branches: (1) a glance branch that coarsely aligns the query and moment proposals using inter-video contrastive learning, and (2) a gaze branch that finely aligns two modalities using both inter- and intra-video contrastive learning. Based on the glance-to-gaze design, our JSG method learns two separate joint embedding spaces for moments and text queries using a hybrid synergistic contrastive learning strategy. Extensive experiments on three public benchmarks, i.e., Charades-STA, DiDeMo, and ActivityNet-Captions demonstrate the superior performance of our JSG method on both video-level retrieval and event-level retrieval subtasks. Our open-source implementation code is available at https://github.com/CFM-MSG/Code_JSG.
Xun Jiang 0001, Xing Xu 0001, Zuo Cao, Yijun Mo, Heng Tao Shen
ACM Multimedia4
2022 Semi-supervised Video Paragraph Grounding with Contrastive Encoder
abstract
Video events grounding aims at retrieving the most relevant moments from an untrimmed video in terms of a given natural language query. Most previous works focus on Video Sentence Grounding (VSG), which localizes the moment with a sentence query. Recently, researchers extended this task to Video Paragraph Grounding (VPG) by retrieving multiple events with a paragraph. However, we find the existing VPG methods may not perform well on context modeling and highly rely on video-paragraph annotations. To tackle this problem, we propose a novel VPG method termed Semi-supervised Video-Paragraph TRansformer (SVPTR), which can more effectively exploit contextual information in paragraphs and significantly reduce the dependency on annotated data. Our SVPTR method consists of two key components: (1) a base model VPTR that learns the video-paragraph alignment with contrastive encoders and tackles the lack of sentence-level contextual interactions and (2) a semi-supervised learning framework with multimodal feature perturbations that reduces the requirements of annotated training data. We evaluate our model on three widely-used video grounding datasets, i.e., ActivityNet-Caption, Charades-CD-OOD, and TACoS. The experimental results show that our SVPTR method establishes the new state-of-the-art performance on all datasets. Even under the conditions of fewer annotations, it can also achieve competitive results compared with recent VPG methods.
Xun Jiang 0001, Xing Xu 0001, Jingran Zhang, Fumin Shen, Zuo Cao, Heng Tao Shen
CVPR5
2022 GTLR: Graph-Based Transformer with Language Reconstruction for Video Paragraph Grounding
abstract
Video Paragraph Grounding aims at retrieving multiple relevant moments from an untrimmed video with a given natural language paragraph query. However, the complex paragraph query brings more challenges to the multimodal fusion and context modeling, which limited the performance of existing VPG methods. To this end, we propose a novel framework for VPG in this paper, termed Graph-based Transformer with Language Reconstruction (GTLR). It consists of three components: (1) Multimodal Graph Encoder conducting the graph reasoning for video-text fusion. (2) Event-wise Decoder predicting the timestamps based on multiple sentence-level features. (3) Language Reconstructor rebuilding the paragraph queries and making our model explainable. We adopt two benchmarks, i.e., ActivityNet-Caption and Charades-STA, to evaluate our model and conduct comprehensive experiments to analyze the effectiveness of each component. The experimental results show that our GTLR method outperforms recent state-of-the-art methods.
Xun Jiang 0001, Xing Xu 0001, Jingran Zhang, Fumin Shen, Zuo Cao
ICME5
2022 Instance-Level Semantic Alignment for Zero-Shot Cross-Modal Retrieval
abstract
Zero-shot Cross-Modal Retrieval (ZS-CMR) is challenging due to the heterogeneous distributions across different modalities and the inconsistent semantics across seen and unseen classes. Previous methods usually perform class-level semantic alignment of data from different modalities by introducing auxiliary word embeddings of class labels, which have a fatal limitation as the learning of class-level information will lead to the ignorance of intra-modal variance. To solve this problem, we propose our Instance-Level Semantic Alignment (ILSA) method to make full use of the instance-level information. We use two disentanglement variational auto-encoders to decompose the data from two modalities into modal specific and modal invariant features. With an instance-level semantic features extractor and a distribution generator, ILSA could generate more appropriate distributions by the learned instance-level semantic features, without any auxiliary knowledge. We perform the experiment on six widely used datasets on two scenarios of ZS-CMR, the results show that our method establishes the new state-of-the-art performance on all datasets.
Yifan Wang 0027, Xing Xu 0001, Zuo Cao
ICME4
2022 Multimodal Disentanglement Variational AutoEncoders for Zero-Shot Cross-Modal Retrieval
abstract
Zero-Shot Cross-Modal Retrieval (ZS-CMR) has recently drawn increasing attention as it focuses on a practical retrieval scenario, i.e., the multimodal test set consists of unseen classes that are disjoint with seen classes in the training set. The recently proposed methods typically adopt the generative model as the main framework to learn a joint latent embedding space to alleviate the modality gap. Generally, these methods largely rely on auxiliary semantic embeddings for knowledge transfer across classes and unconsciously neglect the effect of the data reconstruction manner in the adopted generative model. To address this issue, we propose a novel ZS-CMR model termed Multimodal Disentanglement Variational AutoEncoders (MDVAE), which consists of two coupled disentanglement variational autoencoders (DVAEs) and a fusion-exchange VAE (FVAE). Specifically, DVAE is developed to disentangle the original representations of each modality into modality-invariant and modality-specific features. FVAE is designed to fuse and exchange information of multimodal data by the reconstruction and alignment process without pre-extracted semantic embeddings. Moreover, an advanced counter-intuitive cross-reconstruction scheme is further proposed to enhance the informativeness and generalizability of the modality-invariant features for more effective knowledge transfer. The comprehensive experiments on four image-text retrieval and two image-sketch retrieval datasets consistently demonstrate that our method establishes the new state-of-the-art performance.
Xing Xu 0001, Zuo Cao, Fumin Shen, Heng Tao Shen
SIGIR4
2022 Structure-Aware Semantic-Aligned Network for Universal Cross-Domain Retrieval
abstract
The goal of cross-domain retrieval (CDR) is to search for instances of the same category in one domain by using a query from another domain. Existing CDR approaches mainly consider the standard scenario that the cross-domain data for both training and testing come from the same categories and underlying distributions. However, these methods cannot be well extended to the newly emerging task of universal cross-domain retrieval (UCDR), where the testing data belong to the domain and categories not present during training. Compared to CDR, the UCDR task is more challenging due to (1) visually diverse data from multi-source domains, (2) the domain shift between seen and unseen domains, and (3) the semantic shift across seen and unseen categories. To tackle these problems, we propose a novel model termed Structure-Aware Semantic-Aligned Network (SASA) to align the heterogeneous representations of multi-source domains without loss of generalizability for the UCDR task. Specifically, we leverage the advanced Vision Transformer (ViT) as the backbone and devise a distillation-alignment ViT (DAViT) with a novel token-based strategy, which incorporates two complementary distillation and alignment tokens into the ViT architecture. In addition, the distillation token is devised to improve the generalizability of our model by structure information preservation and the alignment token is used to improve discriminativeness with trainable categorical prototypes. Extensive experiments on three large-scale benchmarks, i.e., Sketchy, TU-Berlin, and DomainNet, demonstrate the superiority of our SASA method over the state-of-the-art UCDR and ZS-SBIR methods.
Xing Xu 0001, Zuo Cao, Heng Tao Shen
SIGIR4
2021 Partial Feature Selection and Alignment for Multi-Source Domain Adaptation
abstract
Multi-Source Domain Adaptation (MSDA), which dedicates to transfer the knowledge learned from multiple source domains to an unlabeled target domain, has drawn increasing attention in the research community. By assuming that the source and target domains share consistent key feature representations and identical label space, existing studies on MSDA typically utilize the entire union set of features from both the source and target domains to obtain the feature map and align the map for each category and domain. However, the default setting of MSDA may neglect the issue of "partialness", i.e., 1) a part of the features contained in the union set of multiple source domains may not present in the target domain; 2) the label space of the target domain may not completely overlap with the multiple source domains. In this paper, we unify the above two cases to a more generalized MSDA task as Multi-Source Partial Domain Adaptation (MSPDA). We propose a novel model termed Partial Feature Selection and Alignment (PFSA) to jointly cope with both MSDA and MSPDA tasks. Specifically, we firstly employ a feature selection vector based on the correlation among the features of multiple sources and target domains. We then design three effective feature alignment losses to jointly align the selected features by preserving the domain information of the data sample clusters in the same category and the discrimination between different classes. Extensive experiments on various benchmark datasets for both MSDA and MSPDA tasks demonstrate that our proposed PFSA approach remarkably outperforms the state-of-the-art MSDA and unimodal PDA methods.
Yangye Fu, Xing Xu 0001, Zuo Cao, Yanli Ji, Kai Zuo, Huimin Lu 0001
CVPR4
2021 Combine Early and Late Fusion Together: A Hybrid Fusion Framework for Image-Text Matching
abstract
Image-text matching is a challenging task in cross-modal learning due to the discrepancy of data representation be-tween different modalities of images and texts. The main-stream methods adopt the late fusion to generate image-text similarity on encoded cross-modal features, and put effort to capture intra-modality associations with considerably high training cost. In this work, we propose to Combine Early and Late Fusion Together (CELFT), which is a universal hybrid fusion framework that can effectively overcome the above shortcomings of the late fusion scheme. In the pro-posed CELFT framework, the hybrid structure with early fusion and late fusion could facilitate the interaction between image and text modalities at early stage. Moreover, these two kinds of fusion strategies complement each other in capturing the inter-modal and intra-modal information, which ensure to learn more accurate image-text similarity. In the experiments, we choose four latest approaches based on the late fusion scheme as the base models, and integrate them with our CELFT framework. The results on two widely used image-text datasets MSCOCO and Flickr30K show that the matching performance of all base models is significantly improved with remarkably reduced training time.
Yifan Wang 0027, Xing Xu 0001, Ruicong Xu, Zuo Cao, Heng Tao Shen
ICME5
2021 Hybrid Fusion with Intra- and Cross-Modality Attention for Image-Recipe Retrieval
abstract
Image-recipe retrieval, which aims at retrieving the relevant recipe from a food image and vice versa, is now attracting widespread attention, since sharing food-related images and recipes on the Internet has become a popular trend. Existing methods have formulated this problem as a typical cross-modal retrieval task by learning the image-recipe similarity. Though these methods have made inspiring achievements for image-recipe retrieval, they may still be less effective to jointly incorporate the three crucial points: (1) the association between ingredients and instructions, (2) fine-grained image information, and (3) the latent alignment between recipes and images. To this end, we propose a novel framework namedHybrid Fusion with Intra- and Cross-Modality Attention (HF-ICMA) to learn accurate image-recipe similarity. Our HF-ICMA model adopts an intra-recipe fusion module to focus on the interaction between ingredients and instructions within a recipe, and further enriches the expressions of the two separate embeddings. Meanwhile, an image-recipe fusion module is devised to explore the potential relationship between fine-grained image regions and ingredients from the recipe, which jointly forms the final image-recipe similarity from both the local and global aspects. Extensive experiments on the large-scale benchmark dataset Recipe1M show that our model significantly outperforms the state-of-the-art approaches on various image-recipe retrieval scenarios.
Xing Xu 0001, Fumin Shen, Zuo Cao, Kai Zuo, Heng Tao Shen
SIGIR5