Yukun Zuo

dblp:229/4037 · DBLP profile ↗
← Back
9ranked-venue papers
6as first author
8since 2021 · last 2025
0000-0002-2883-8572ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Resource-Efficient Affordance Grounding with Complementary Depth and Semantic Prompts
abstract
Affordance refers to the functional properties that an agent perceives and utilizes from its environment, and is key perceptual information required for robots to perform actions. This information is rich and multimodal in nature. Existing multimodal affordance methods face limitations in extracting useful information, mainly due to simple structural designs, basic fusion methods, and large model parameters, making it difficult to meet the performance requirements for practical deployment. To address these issues, this paper proposes the BiT-Align image-depth-text affordance mapping framework. The framework includes a Bypass Prompt Module (BPM) and a Text Feature Guidance (TFG) attention selection mechanism. BPM integrates the auxiliary modality depth image directly as a prompt to the primary modality RGB image, embedding it into the primary modality encoder without introducing additional encoders. This reduces the model’s parameter count and effectively improves functional region localization accuracy. The TFG mechanism guides the selection and enhancement of attention heads in the image encoder using textual features, improving the understanding of affordance characteristics. Experimental results demonstrate that the proposed method achieves significant performance improvements on public AGD20K and HICO-IIF datasets. On the AGD20K dataset, compared with the current state-of-the-art method, we achieve a 6.0% improvement in the KLD metric, while reducing model parameters by 88.8%, demonstrating practical application values. The source code will be made publicly available at https://github.com/DAWDSE/BiT-Align.
Fan Yang 0063, Guoliang Zhu, Hao Shi 0004, Yukun Zuo, Wenrui Chen, Zhiyong Li 0001, Kailun Yang 0001
IROS6
2025 Exploring Communication and Roadside Perception Requirements for Cooperative Warning Systems at Intersections
abstract
Infrastructure-based cooperative perception has been researched for several years, but few automotive warning or control applications using this information have been published. Infrastructure sensing, such as with cameras or lidars, and a communication system, allows connected vehicles to receive information about all observed objects. An SAE standard, “V2X Sensor-Sharing for Cooperative and Automated Driving” (J3224), released in 2022, introduces the Sensor Data Sharing Message (SDSM) as the standard communication message for cooperative perception. This paper investigates the use of the SDSM for a vehicle application to provide warnings of potential collisions with vulnerable road users who will cross the street at the intersection. The application was tested in CARLA simulation under various roadside detection errors and communication conditions to assess the impact on the on-board application and estimate the minimum detection and communication requirements for effective use. In addition, the system was implemented and evaluated at the Mcity test facility. The results demonstrate that the proposed warning system can accurately and promptly warn the driver, given specific communication conditions, and show that the SDSM is viable for real-time on-board usage.
Tinghan Wang, Depu Meng, Boqi Li 0001, Rusheng Zhang, Yukun Zuo, Shengyin Shen, Darian Hogue, Michael Maile, Michael Shulman, Henry X. Liu
IV5
2024 Hierarchical Augmentation and Distillation for Class Incremental Audio-Visual Video Recognition
abstract
Audio-visual video recognition (AVVR) integrates audio and visual cues to accurately categorize videos. While current methods using provided datasets achieve satisfactory results, they face challenges in retaining historical class knowledge when new classes appear in real-world situations. There are no dedicated methods to address this issue, prompting this paper to explore Class Incremental Audio-Visual Video Recognition (CIAVVR). CIAVVR aims to preserve historical knowledge contained in stored data and learned models to prevent catastrophic forgetting. Audio-visual data and models inherently have hierarchical structures, where the model contains both low-level and high-level semantic information, and data includes snippet-level, video-level, and distribution-level spatial information. It is crucial to fully exploit these hierarchical structures for data knowledge preservation and model knowledge preservation. However, existing image class incremental learning methods do not explicitly consider these hierarchical structures. Therefore, we introduce Hierarchical Augmentation and Distillation (HAD), which includes the Hierarchical Augmentation Module (HAM) and Hierarchical Distillation Module (HDM). These modules efficiently utilize the hierarchical structure of data and models. Specifically, HAM uses a novel augmentation strategy, segmental feature augmentation, to preserve hierarchical model knowledge. Simultaneously, HDM employs newly designed hierarchical logical distillation (video-distribution) and hierarchical correlative distillation (snippet-video) to maintain intra-sample and inter-sample hierarchical knowledge. Evaluations on four benchmarks (AVE, AVK-100, AVK-200, and AVK-400) show that HAD effectively captures hierarchical information, enhancing the preservation of historical class knowledge and performance. We also provide a theoretical analysis to support the segmental feature augmentation strategy.
Yukun Zuo, Hantao Yao, Liansheng Zhuang, Changsheng Xu
IEEE Trans. Pattern Anal. Mach. Intell.1
2024 Incremental Audio-Visual Fusion for Person Recognition in Earthquake Scene
abstract
Earthquakes have a profound impact on social harmony and property, resulting in damage to buildings and infrastructure. Effective earthquake rescue efforts require rapid and accurate determination of whether any survivors are trapped in the rubble of collapsed buildings. While deep learning algorithms can enhance the speed of rescue operations using single-modal data (either visual or audio), they are confronted with two primary challenges: insufficient information provided by single-modal data and catastrophic forgetting. In particular, the complexity of earthquake scenes means that single-modal features may not provide adequate information. Additionally, catastrophic forgetting occurs when the model loses the information learned in a previous task after training on subsequent tasks, due to non-stationary data distributions in changing earthquake scenes. To address these challenges, we propose an innovative approach that utilizes an incremental audio-visual fusion model for person recognition in earthquake rescue scenarios. Firstly, we leverage a cross-modal hybrid attention network to capture discriminative temporal context embedding, which uses self-attention and cross-modal attention mechanisms to combine multi-modality information, enhancing the accuracy and reliability of person recognition. Secondly, an incremental learning model is proposed to overcome catastrophic forgetting, which includes elastic weight consolidation and feature replay modules. Specifically, the elastic weight consolidation module slows down learning on certain weights based on their importance to previously learned tasks. The feature replay module reviews the learned knowledge by reusing the features conserved from the previous task, thus preventing catastrophic forgetting in dynamic environments. To validate the proposed algorithm, we collected the Audio-Visual Earthquake Person Recognition (AVEPR) dataset from earthquake films and real scenes. Furthermore, the proposed method gets 85.41% accuracy while learning the 10th new task, which demonstrates the effectiveness of the proposed method and highlights its potential to significantly improve earthquake rescue efforts.
Sisi You, Yukun Zuo, Hantao Yao, Changsheng Xu
ACM Trans. Multim. Comput. Commun. Appl.2
2023 Dual Structural Knowledge Interaction for Domain Adaptation
abstract
Domain adaptation aims to transfer knowledge from a label-rich source domain to an unlabeled target domain. A common strategy is to assign pseudo-labels to unlabeled target samples for performing representation learning. However, most existing methods only apply the source-guided classifier to generate the source-biased pseudo-labels for self-training, leading to biased target representations. Moreover, the generated pseudo-labels ignore the manifold assumption that neighboring samples are likely to have the same labels. To address the above problem, we formulate a novel structural knowledge to assign target-oriented and manifold-guided pseudo-labels for unlabeled target samples. The structural knowledge consists of cluster-based knowledge and locality-based knowledge. The cluster-based knowledge denotes the label consistency between the target samples and the non-parametric target cluster centers, making the pseudo-labels target-oriented. The locality-based knowledge constrains the target sample and its neighbors to satisfy the manifold assumption. As the neighbors contain the source and target samples, the source and target locality-based knowledge are utilized to boost the descriptions. With the structural knowledge, we propose a novel Dual Structural Knowledge Interaction (DSKI) framework for domain adaptation. For generating aligned and discriminative features, knowledge consistency constraint and instance mutual constraint are proposed in DSKI. Evaluations on three benchmarks demonstrate the effectiveness of the Dual Structural Knowledge Interaction,e.g.,74.9%, 87.7%, and 90.8% for Office-Home, VisDa-2017, and Office-31, respectively.
Yukun Zuo, Hantao Yao, Liansheng Zhuang, Changsheng Xu
IEEE Trans. Multim.1
2022 Margin-Based Adversarial Joint Alignment Domain Adaptation
abstract
Domain adaptation aims to transfer the knowledge learned from a labeled source domain to an unlabeled target domain, which has different data distribution with the source domain. Most of the existing methods focus on aligning the data distribution between the source and target domains but ignore the discrimination of the feature space among categories, leading the samples close to the decision boundary to be misclassified easily. To address the above issue, we propose a Margin-based Adversarial Joint Alignment (MAJA) to constrain the feature spaces of source and target domains to be aligned and discriminative. The proposed MAJA consists of two components: joint alignment module and margin-based generative module. The joint alignment module is proposed to align the source and target feature spaces by considering the joint distribution of features and labels. Therefore, the embedding features and the corresponding labels treated as pair data are applied for domain alignment. Furthermore, the margin-based generative module is proposed to boost the discrimination of the feature space,i.e.,make all samples as far away from the decision boundary as possible. The margin-based generative module first employs the Generative Adversarial Networks (GAN) to generate a lot of fake images for each category, then applies the adversarial learning to enlarge and reduce the category margin for the true images and generated fake images, respectively. The evaluations on three benchmarks,e.g.,small image datasets, VisDA-2017, and Office-31, verify the effectiveness of the proposed method.
Yukun Zuo, Hantao Yao, Liansheng Zhuang, Changsheng Xu
IEEE Trans. Circuits Syst. Video Technol.1
2022 Seek Common Ground While Reserving Differences: A Model-Agnostic Module for Noisy Domain Adaptation
abstract
Noisy domain adaptation aims to solve the problem that the source dataset contains noisy labels in domain adaptation. Previous methods handle noisy labels by selecting the small-loss samples with inconsistent predictions between two models and discarding the consistent samples, resulting in many noises contained in the selected samples. By jointly considering the consistent and inconsistent samples, we propose a model-agnostic module, named Seek Common Ground While Reserving Differences (SCGWRD), to reduce the impact of noisy samples. The proposed SCGWRD module consists of Seek Common Ground (SCG) component and Reserve Differences (RD) component by utilizing the outputs of two symmetrical domain adaptation models. As the common samples with consistent predictions between two models are more likely to be clean samples, the SCG component applies the small-loss strategy to select the reliable samples with consistent predictions. Unlike SCG, the RD component maintains the divergences between two models with mutual learning and reduces the effect of noisy data using the samples with different predictions and small losses. Evaluations on three benchmarks demonstrate the effectiveness and robustness of the proposed SCGWRD module for noisy domain adaptation.
Yukun Zuo, Hantao Yao, Liansheng Zhuang, Changsheng Xu
IEEE Trans. Multim.1
2021 Attention-Based Multi-Source Domain Adaptation
abstract
Multi-source domain adaptation (MSDA) aims to transfer knowledge from multi-source domains to one target domain. Inspired by single-source domain adaptation, existing methods solve MSDA by aligning the data distributions between the target domain and each source domain. However, aligning the target domain with the dissimilar source domain would harm the representation learning. To address the above issue, an intuitive motivation of MSDA is using the attention mechanism to enhance the positive effects of the similar domains, and suppress the negative effects of the dissimilar domains. Therefore, we propose Attention-Based Multi-Source Domain Adaptation (ABMSDA) by considering the domain correlations to alleviate the effects caused by dissimilar domains. To obtain the domain correlations between source and target domains, ABMSDA firstly trains a domain recognition model to calculate the probability that the target images belong to each source domain. Based on the domain correlations, Weighted Moment Distance (WMD) is proposed to pay more attention on the source domains with higher similarities. Furthermore, Attentive Classification Loss (ACL) is developed to constrain that the feature extractor can generate the alignment and discriminative visual representations. The evaluations on two benchmarks demonstrate the effectiveness of the proposed model, e.g., an average of 6.1% improvement on the challenging DomainNet dataset.
Yukun Zuo, Hantao Yao, Changsheng Xu
IEEE Trans. Image Process.1
2020 Category-Level Adversarial Self-Ensembling for Domain Adaptation
abstract
Domain adaptation aims at learning from a source data distribution a well-performing model on a different target data distribution. Recently, the self-ensembling-based methods have been proved to be effective for unsupervised domain adaptation. However, they still have two shortcomings: 1) no explicitly constraint about the distributions between the source and target domains; 2) the Euclidean distance fails to measure the similarity between two distributions with no overlap. To solve those shortcomings, we propose a novel Category-level Adversarial Self-ensembling (CAS) model for domain adaptation, which contains two types of consistency constraints. The first one is how to constrain the descriptions for source and target domains to be aligned. Therefore, we adopt a minimax game with a discrepancy loss between the category information generated by two classifiers. As the self-ensembling consists of two sub-networks: student and teacher networks, the second one is the consistency between those two networks for the target samples. Aiming to overcome the disadvantage of the Euclidean metric, we employ the Wasserstein distance to measure the difference between two probabilistic distributions. Experiments on several benchmarks demonstrate that our proposed CAS is superior to existing methods.
Yukun Zuo, Hantao Yao, Changsheng Xu
ICME1