Yaguang Song

dblp:249/8313 · DBLP profile ↗
← Back
13ranked-venue papers
4as first author
12since 2021 · last 2026
0000-0002-9300-8110ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 3 first-author · 10 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 4 since 2021Computer networks · 3 · 2 first-author · 3 since 2021
YearPublicationVenuePosition
2026 Sample-Aware Knowledge Association and Enhancement for Open-Vocabulary Continual Learning
Zhilin Zhu 0001, Zhiheng Ma, Yabin Wang 0001, Yaguang Song, Yaowei Wang 0001, Xiaopeng Hong
Int. J. Comput. Vis.4
2025 Pilot: Building the Federated Multimodal Instruction Tuning Framework
abstract
In this paper, we explore a novel federated multimodal instruction tuning task(FedMIT), which is significant for collaboratively fine-tuning MLLMs on different types of multimodal instruction data on distributed devices. To solve the new task, we propose a federated multimodal instruction tuning framework(Pilot). Our framework integrates two-stage of ``adapter on adapter” into the connector of the vision encoder and the LLM. In stage 1, we extract task-specific features and client-specific features from visual information. In stage 2, we build the cross-task Mixture-of-Adapters(CT-MoA) module to perform cross-task interaction. Each client can not only capture personalized information of local data and learn task-related multimodal information, but also learn general knowledge from other tasks. In addition, we introduce an adaptive parameter aggregation strategy for text training parameters, which optimizes parameter aggregation by calculating weights based on the euclidean distance between parameters, so that parameter aggregation can benefit from positive effects to the greatest extent while effectively reducing negative effects. Our framework can collaboratively exploit distributed data from different local clients to learn cross-task knowledge without being affected by the task heterogeneity during instruction tuning. The effectiveness of our method is verified in two different cross-task scenarios.
Baochen Xiong, Xiaoshan Yang, Yaguang Song, Yaowei Wang 0001, Changsheng Xu
AAAI3
2025 Towards A Real-World Road Damage Detection Dataset
abstract
Road damage represents a serious challenge to the health of road infrastructure and driving safety, making deep learning-based image analysis for road damage detection (RDD) an important research focus. The limited diversity in road damage types, road image collection, size, environment, and imperfect damage definitions within current RDD datasets restrict the real-world applications of RDD. To address this issue, this paper constructs PCL-RDD, a new and extensive RDD dataset. It comprises 24,765 road images, 54,732 instances, and 19 types of road damage. Compared to the existing datasets that mainly include common road damage, the proposed dataset contains a variety of rare and urgent road damages. Besides, we collect road facility-related damages, which also affect traffic safety. We evaluate eight well-established object detection algorithms on the dataset, highlighting the limitations of state-of-the-art detection algorithms under complex conditions. This study contributes a significant dataset to the RDD field and can advance artificial intelligence in both city infrastructure management and environmental perception for autonomous driving. The dataset is available at https://github.com/humh-c/PCL-RDD.
Menghao Hu, Zuogan Tang, Xiaoshan Yang, Zhe Wu 0006, Zhouxin Yang, Shaocong Wu, Yaguang Song, Kui Hou, Yaowei Wang 0001
ICME7
2025 Health-oriented Multimodal Food Question Answering with Implicit and Explicit Knowledge
abstract
Health-oriented food analysis has become a research hotspot in recent years because it can help people keep away from unhealthy diets. Remarkable advancements have been made in recipe retrieval, food recommendation, nutrition analysis, and calorie estimation. However, existing works still cannot well balance the individual preference and the health. Multimodal food question and answering (MFQA) presents substantial promise for practical applications, yet it remains underexplored. In this article, we introduce a health-oriented MFQA dataset with 9,000 Chinese question−answer pairs based on a multimodal food knowledge graph (MFKG) collected from a food-sharing Web site. Additionally, we propose a novel framework for MFQA in the health domain that leverages implicit general knowledge and explicit domain-specific knowledge. The framework comprises four key components: implicit general knowledge injection module (IGKIM), explicit domain-specific knowledge retrieval module (EDKRM), ranking module, and answer module. The IGKIM facilitates knowledge acquisition at both the feature and text levels. The EDKRM retrieves the most relevant candidate knowledge from the knowledge graph based on the given question. The ranking module sorts the results retrieved by EDKRM and further retrieve candidate knowledge relevant to the problem. Subsequently, the answer module thoroughly analyzes the multimodal information in the query along with the retrieved relevant knowledge to predict accurate answers. Extensive experimental results on the MFQA dataset demonstrate the effectiveness of our proposed method. The code and dataset are available at https://github.com/Wjianghai/HMFQA .
Menghao Hu, Yaguang Song, Xiaoshan Yang, Yaowei Wang 0001, Changsheng Xu
ACM Trans. Multim. Comput. Commun. Appl.2
2024 Modality-Collaborative Test-Time Adaptation for Action Recognition
abstract
Video-based Unsupervised Domain Adaptation (VUDA) method improves the generalization of the video model, en-abling it to be applied to action recognition tasks in different environments. However, these methods require contin-uous access to source data during the adaptation process, which are impractical in real scenarios where the source videos are not available with concerns in transmission efficiency or privacy issues. To address this problem, in this paper, we focus on the Multimodal Video Test- Time Adaptation (MVTTA) task. Existing image-based TTA methods cannot be directly applied to this task because videos have domain shifts in multimodal and temporal, which brings difficulties to adaptation. To address the above challenges, we propose a Modality-Collaborative Test-Time Adaptation (MC-TTA) Network. MC-TTA contains maintain teacher and student memory banks respectively for generating pseudo-prototypes and target-prototypes. In the teacher model, we propose Self-assembled Source-friendly Feature Reconstruction (SSFR) to encourage the teacher memory bank to store features that are more likely to be consistent with the source distribution. Through multimodal prototype alignment and cross-modal relative consistency, our method can effectively alleviate domain shift in videos. We evaluate the proposed model on four public video datasets. The results show that our model outperforms existing state-of-the-art methods.
Baochen Xiong, Xiaoshan Yang, Yaguang Song, Yaowei Wang 0001, Changsheng Xu
CVPR3
2024 Libra: Building Decoupled Vision System on Large Language Models
abstract
In this work, we introduce **Libra**, a prototype model with a decoupled vision system on a large language model (LLM). The decoupled vision system decouples inner-modal modeling and cross-modal interaction, yielding unique visual information modeling and effective cross-modal comprehension. Libra is trained through discrete auto-regressive modeling on both vision and language inputs. Specifically, we incorporate a routed visual expert with a cross-modal bridge module into a pretrained LLM to route the vision and language flows during attention computing to enable different attention patterns in inner-modal modeling and cross-modal interaction scenarios. Experimental results demonstrate that the dedicated design of Libra achieves a strong MLLM baseline that rivals existing works in the image-to-text scenario with merely 50 million training data, providing a new perspective for future multimodal foundation models. Code is available at https://github.com/YifanXu74/Libra.
Yifan Xu 0008, Xiaoshan Yang, Yaguang Song, Changsheng Xu
ICML3
2024 Part-Aware Prompt Tuning for Weakly Supervised Referring Expression Grounding
Chenlin Zhao, Jiabo Ye, Yaguang Song, Ming Yan 0008, Xiaoshan Yang, Changsheng Xu
MMM (3)3
2024 Recovering Generalization via Pre-Training-Like Knowledge Distillation for Out-of-Distribution Visual Question Answering
abstract
With the emergence of large-scale multi-modal foundation models, significant improvements have been made towards Visual Question Answering (VQA) in recent years via the “Pre-training and Fine-tuning” paradigm. However, the fine-tuned VQA model, which is more specialized for the downstream training data, may fail to generalize well when there is a distribution shift between the training and test data, which is defined as the Out-of-Distribution (OOD) problem. An intuitive way to solve this problem is to transfer the common knowledge from the foundation model to the fine-tuned VQA model via knowledge distillation for better generalization. However, the generality of distilled knowledge based on the task-specific training data is questionable due to the bias between the training and test data. An ideal way is to adopt the pre-training data to distill the common knowledge shared by the training and OOD test samples, which however is impracticable due to the huge size of pre-training data. Based on the above considerations, in this article, we propose a method, named Pre-training-like Knowledge Distillation (PKD), to imitate the pre-training feature distribution and leverage it to distill the common knowledge, which can improve the generalization performance of the fine-tuned model for OOD VQA. Specifically, we first leverage the in-domain VQA data as guidance and adopt two cross-modal feature prediction networks, which are learned under the supervision of image-text matching loss and feature divergence loss, to estimate pre-training-like vision and text features. Next, we conduct feature-level distillation by explicitly integrating the downstream VQA input features with the predicted pre-training-like features through a memory mechanism. In the meantime, we also conduct model-level distillation by constraining the image-text matching output of the downstream VQA model and the output of the foundation model for the pre-training-like image and text features. Extensive experiments on the VQA-CP v2 and VQA v2 datasets demonstrate the effectiveness of our method.
Yaguang Song, Xiaoshan Yang, Yaowei Wang 0001, Changsheng Xu
IEEE Trans. Multim.1
2023 Client-Adaptive Cross-Model Reconstruction Network for Modality-Incomplete Multimodal Federated Learning
abstract
Multimodal federated learning (MFL) is an emerging field that allows many distributed clients, each with multimodal data, to work together to train models targeting multimodal tasks without sharing local data. Whereas, existing methods assume that all modalities for each sample are complete, which limits their practicality. In this paper, we propose a Client-Adaptive Cross-Modal Reconstruction Network (CACMRN) to solve the modality-incomplete multimodal federated learning (MI-MFL). Compared to existing centralized methods for reconstructing missing modality, the local client data in federated learning is typically much less, which makes it challenging to train a reliable reconstruction model that can accurately predict missing data. We propose a cross-modal reconstruction transformer, which can prevent the model overfitting on the local client by exploring instance-instance relationships within the local client and utilizing normalized self-attention to conduct data-depended partial updating. Using federated optimization with alternative local updating and global aggregation, our method can not only collaboratively utilize the distributed data on different local clients to learn the cross-modal reconstruction transformer, but also prevent the reconstruction model from overfitting the data on the local client. Extensive experimental results on three datasets demonstrate the effectiveness of our method.
Baochen Xiong, Xiaoshan Yang, Yaguang Song, Yaowei Wang 0001, Changsheng Xu
ACM Multimedia3
2023 Health-Oriented Multimodal Food Question Answering
Jianghai Wang, Menghao Hu, Yaguang Song, Xiaoshan Yang
MMM (1)3
2023 Self-supervised Calorie-aware Heterogeneous Graph Networks for Food Recommendation
abstract
With the rapid development of online recipe sharing platforms, food recommendation is emerging as an important application. Although recent studies have made great progress on food recommendation, they have two shortcomings that are likely to affect the recommendation performance. (1) The relations between ingredients are not considered, which may lead to sub-optimal representations of recipes and further result in the neglect of the user’s personalized ingredient combination preference. (2) Existing methods do not consider the impact of users’ preferences on calories in users’ food decision-making process. In this article, we propose a Self-supervised Calorie-aware Heterogeneous Graph Network (SCHGN) to model the relations between ingredients and incorporate calories of food simultaneously. Specifically, we first incorporate users, recipes, ingredients, and calories into a heterogeneous graph and explicitly present the complex relations among them with directed edges. Then, we explore the co-occurrence relation of ingredients in different recipes via self-supervised ingredient prediction. To capture users’ dynamic preferences on calories of food, we learn calorie-aware user representations by hierarchical message passing and compute a comprehensive user-guided recipe representation by attention mechanism. The final food recommendation is accomplished based on the similarity between a user’s calorie-aware representation and the user-guided representation of a recipe. Extensive experiment results on benchmark datasets demonstrate the effectiveness of the proposed method.
Yaguang Song, Xiaoshan Yang, Changsheng Xu
ACM Trans. Multim. Comput. Commun. Appl.1
2022 Learning Hierarchical Video Graph Networks for One-Stop Video Delivery
abstract
The explosive growth of video data has brought great challenges to video retrieval, which aims to find out related videos from a video collection. Most users are usually not interested in all the content of retrieved videos but have a more fine-grained need. In the meantime, most existing methods can only return a ranked list of retrieved videos lacking a proper way to present the video content. In this paper, we introduce a distinctively new task, namely One-Stop Video Delivery (OSVD) aiming to realize a comprehensive retrieval system with the following merits: it not only retrieves the relevant videos but also filters out irrelevant information and presents compact video content to users, given a natural language query and video collection. To solve this task, we propose an end-to-end Hierarchical Video Graph Reasoning framework (HVGR) , which considers relations of different video levels and jointly accomplishes the one-stop delivery task. Specifically, we decompose the video into three levels, namely the video-level, moment-level, and the clip-level in a coarse-to-fine manner, and apply Graph Neural Networks (GNNs) on the hierarchical graph to model the relations. Furthermore, a pairwise ranking loss named Progressively Refined Loss is proposed based on prior knowledge that there is a relative order of the similarity of query-video, query-moment, and query-clip due to the different granularity of matched information. Extensive experimental results on benchmark datasets demonstrate that the proposed method achieves superior performance compared with baseline methods.
Yaguang Song, Junyu Gao 0002, Xiaoshan Yang, Changsheng Xu
ACM Trans. Multim. Comput. Commun. Appl.1
2019 EEG-Based Motor Imagery Classification with Deep Multi-Task Learning
abstract
In the past decade, Electroencephalogram (EEG) has been applied in many fields, such as Motor Imagery (MI) and Emotion Recognition. Traditionally, for classification tasks based on EEG, researchers would extract features from raw signals manually which is often time consuming and requires adequate domain knowledge. Besides that, features manually extracted and selected may not generalize well due to the limitation of human. Convolutional Neural Networks (CNNs) plays an important role in the wave of deep learning and achieve amazing results in many areas. One of the most attractive features of deep learning for EEG-based tasks is the end-to-end learning. Features are learned from raw signals automatically and the feature extractor and classifier are optimized simultaneously. There are some researchers applying deep learning methods to EEG analysis and achieving promising performances. However, supervised deep learning methods often require large-scale annotated dataset, which is almost impossible to acquire in EEG-based tasks. This problem limits the further improvements of deep learning models for classification based on EEG. In this paper, we propose a novel deep learning method DMTL-BCI based on Multi-Task Learning framework for EEG-based classification tasks. The proposed model consists of three modules, the representation module, the reconstruction module and the classification module. Our model is proposed to improve the classification performance with limited EEG data. Experimental results on benchmark dataset, BCI Competition IV dataset 2a, show that our proposed method outperforms the state-of-the-art method by 3.0%, which demonstrates the effectiveness of our model.
Yaguang Song, Danli Wang, Kang Yue, Zuo-Jun Max Shen
IJCNN1