Yun Liu 0017

dblp:50/2482-17 · DBLP profile ↗
← Back
17ranked-venue papers
10as first author
14since 2021 · last 2026
0000-0002-3557-6179ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 9 first-author · 9 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Multimodal sentiment analysis based on label semantic guidance under social links
Yun Liu 0017, Guofeng He, Zhoujun Li 0001
Pattern Recognit.1
2025 Learning fine-grained representation with token-level alignment for multimodal sentiment analysis
Xiang Li 0117, Haijun Zhang 0007, Zhiqiang Dong, Xianfu Cheng, Yun Liu 0017, Xiaoming Zhang 0001
Expert Syst. Appl.5
2025 Moment matching of joint distributions for unsupervised domain adaptation
Bo Zhang 0096, Xiaoming Zhang 0001, Yun Liu 0017, Yancong Li, Feiran Huang
Inf. Process. Manag.4
2025 Vision-language representation learning with breadth and depth attention pre-training
abstract
The rapid advances in computer vision and natural language processing have led to increased attention toward the challenge of understanding vision and language together across multiple domains. Representation learning has become a major focus of research on cross-modal information understanding. However, current methods often fall short of providing comprehensive interaction and meaningful supervised guidance that would allow for effective learning of visual-linguistic joint representation. In this paper, we introduce the Breadth and Depth Attention Pre-training (BDAP) model for vision–language representation learning. Our model includes a breadth attention network designed to model feature associations between text sentences and image regions across different image levels. It uses fine-grained image features to promote more effective cross-modal feature interactions. Additionally, a depth attention network, which repeatedly calculates attention scores , is designed to deeply capture the complementarity between the image and text by gradually refining important image regions related to the text. Furthermore, we propose an attention pre-training network that leverages attention annotated distribution maps as prior knowledge to supervise the learning process of the breadth and depth attention networks, thereby enabling weight initialization of both types of attention networks. Extensive experiments on datasets of visual question answering and multi-modal sentiment analysis demonstrate the promising superiority of our BDAP model for vision–language representation learning.
Yun Liu 0017, Bo Zhang 0096, Chencheng Wang, Genglong Yan, Zhoujun Li 0001
Knowl. Based Syst.1
2023 Generative Sentiment Transfer via Adaptive Masking
Yingze Xie, Jie Xu 0015, LiQiang Qiao, Yun Liu 0017, Feiran Huang, Chaozhuo Li
PAKDD (4)4
2023 Scanning, attention, and reasoning multimodal content for sentiment analysis
Yun Liu 0017, Zhoujun Li 0001, Shixun Shen
Knowl. Based Syst.1
2022 Dynamic self-attention with vision synchronization networks for video question answering
Yun Liu 0017, Xiaoming Zhang 0001, Feiran Huang, Shixun Shen, Zhoujun Li 0001
Pattern Recognit.1
2022 ALSA: Adversarial Learning of Supervised Attentions for Visual Question Answering
abstract
Visual question answering (VQA) has gained increasing attention in both natural language processing and computer vision. The attention mechanism plays a crucial role in relating the question to meaningful image regions for answer inference. However, most existing VQA methods: 1) learn the attention distribution either from free-form regions or detection boxes in the image, which is intractable in answering questions about the foreground object and background form, respectively and 2) neglect the prior knowledge of human attention and learn the attention distribution with an unguided strategy. To fully exploit the advantages of attention, the learned attention distribution should focus more on the question-related image regions, such as human attention for both the questions, about the foreground object and background form. To achieve this, this article proposes a novel VQA model, called adversarial learning of supervised attentions (ALSAs). Specifically, two supervised attention modules: 1) free form-based and 2) detection-based, are designed to exploit the prior knowledge for attention distribution learning. To effectively learn the correlations between the question and image from different views, that is, free-form regions and detection boxes, an adversarial learning mechanism is implemented as an interplay between two supervised attention modules. The adversarial learning reinforces the two attention modules mutually to make the learned multiview features more effective for answer inference. The experiments performed on three commonly used VQA datasets confirm the favorable performance of ALSA.
Yun Liu 0017, Xiaoming Zhang 0001, Zhiyun Zhao, Bo Zhang 0096, Zhoujun Li 0001
IEEE Trans. Cybern.1
2022 Cross-Attentional Spatio-Temporal Semantic Graph Networks for Video Question Answering
abstract
Due to the rich spatio-temporal visual content and complex multimodal relations, Video Question Answering (VideoQA) has become a challenging task and attracted increasing attention. Current methods usually leverage visual attention, linguistic attention, or self-attention to uncover latent correlations between video content and question semantics. Although these methods exploit interactive information between different modalities to improve comprehension ability, inter- and intra-modality correlations cannot be effectively integrated in a uniform model. To address this problem, we propose a novel VideoQA model called Cross-Attentional Spatio-Temporal Semantic Graph Networks (CASSG). Specifically, a multi-head multi-hop attention module with diversity and progressivity is first proposed to explore fine-grained interactions between different modalities in a crossing manner. Then, heterogeneous graphs are constructed from the cross-attended video frames, clips, and question words, in which the multi-stream spatio-temporal semantic graphs are designed to synchronously reasoning inter- and intra-modality correlations. Last, the global and local information fusion method is proposed to coalesce the local reasoning vector learned from multi-stream spatio-temporal semantic graphs and the global vector learned from another branch to infer the answer. Experimental results on three public VideoQA datasets confirm the effectiveness and superiority of our model compared with state-of-the-art methods.
Yun Liu 0017, Xiaoming Zhang 0001, Feiran Huang, Bo Zhang 0096, Zhoujun Li 0001
IEEE Trans. Image Process.1
2021 Matching Distributions between Model and Data: Cross-domain Knowledge Distillation for Unsupervised Domain Adaptation
abstract
Bo Zhang, Xiaoming Zhang, Yun Liu, Lei Cheng, Zhoujun Li. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Bo Zhang 0096, Xiaoming Zhang 0001, Yun Liu 0017, Zhoujun Li 0001
ACL/IJCNLP (1)3
2021 Discriminative Feature Adaptation via Conditional Mean Discrepancy for Cross-Domain Text Classification
Bo Zhang 0096, Xiaoming Zhang 0001, Yun Liu 0017
DASFAA (2)3
2021 Dual self-attention with co-attention networks for visual question answering
Yun Liu 0017, Xiaoming Zhang 0001, Qianyun Zhang 0001, Chaozhuo Li, Feiran Huang, Xianghong Tang, Zhoujun Li 0001
Pattern Recognit.1
2021 Adversarial Learning With Multi-Modal Attention for Visual Question Answering
abstract
Visual question answering (VQA) has been proposed as a challenging task and attracted extensive research attention. It aims to learn a joint representation of the question-image pair for answer inference. Most of the existing methods focus on exploring the multi-modal correlation between the question and image to learn the joint representation. However, the answer-related information is not fully captured by these methods, which results that the learned representation is ineffective to reflect the answer of the question. To tackle this problem, we propose a novel model, i.e., adversarial learning with multi-modal attention (ALMA), for VQA. An adversarial learning-based framework is proposed to learn the joint representation to effectively reflect the answer-related information. Specifically, multi-modal attention with the Siamese similarity learning method is designed to build two embedding generators, i.e., question-image embedding and question-answer embedding. Then, adversarial learning is conducted as an interplay between the two embedding generators and an embedding discriminator. The generators have the purpose of generating two modality-invariant representations for the question-image and question-answer pairs, whereas the embedding discriminator aims to discriminate the two representations. Both the multi-modal attention module and the adversarial networks are integrated into an end-to-end unified framework to infer the answer. Experiments performed on three benchmark data sets confirm the favorable performance of ALMA compared with state-of-the-art approaches.
Yun Liu 0017, Xiaoming Zhang 0001, Feiran Huang, Zhoujun Li 0001
IEEE Trans. Neural Networks Learn. Syst.1
2021 Deep Attentive Multimodal Network Representation Learning for Social Media Images
abstract
The analysis for social networks, such as the socially connected Internet of Things, has shown a deep influence of intelligent information processing technology on industrial systems for Smart Cities. The goal of social media representation learning is to learn dense, low-dimensional, and continuous representations for multimodal data within social networks, facilitating many real-world applications. Since social media images are usually accompanied by rich metadata (e.g., textual descriptions, tags, groups, and submitted users), simply modeling the image is not effective to learn the comprehensive information from social media images. In this work, we treat the image and its textual description as multimodal content, and transform other metainformation into the links between contents (such as two images marked by the same tag or submitted by the same user). Based on the multimodal content and social links, we propose a Deep Attentive Multimodal Graph Embedding model named DAMGE for more effective social image representation learning. We introduce both small- and large-scale datasets to conduct extensive experiments, of which the results confirm the superiority of the proposal on the tasks of social image classification and link prediction.
Feiran Huang, Chaozhuo Li, Boyu Gao 0003, Yun Liu 0017, Sattam Al Otaibi, Hao Chen 0062
ACM Trans. Internet Techn.4
2020 Visual Question Answering via Combining Inferential Attention and Semantic Space Mapping
Yun Liu 0017, Xiaoming Zhang 0001, Feiran Huang, Zhonghua Zhao, Zhoujun Li 0001
Knowl. Based Syst.1
2019 Adversarial Learning for Weakly-Supervised Social Network Alignment
abstract
Nowadays, it is common for one natural person to join multiple social networks to enjoy different kinds of services. Linking identical users across multiple social networks, also known as social network alignment, is an important problem of great research challenges. Existing methods usually link social identities on the pairwise sample level, which may lead to undesirable performance when the number of available annotations is limited. Motivated by the isomorphism information, in this paper we consider all the identities in a social network as a whole and perform social network alignment from the distribution level. The insight is that we aim to learn a projection function to not only minimize the distance between the distributions of user identities in two social networks, but also incorporate the available annotations as the learning guidance. We propose three models SNNAu, SNNAb and SNNAo to learn the projection function under the weakly-supervised adversarial learning framework. Empirically, we evaluate the proposed models over multiple datasets, and the results demonstrate the superiority of our proposals.
Chaozhuo Li, Senzhang Wang, Philip S. Yu, Yanbo Liang, Yun Liu 0017, Zhoujun Li 0001
AAAI6
2018 Adversarial Learning of Answer-Related Representation for Visual Question Answering
abstract
Visual Question Answering (VQA) aims to learn a joint embedding of the question sentence and the corresponding image to infer the answer. Existing approaches learn the joint embedding don't consider the answer-related information, which results in that the learned representation is not effective to reflect the answer of the question. To address this problem, this paper proposes a novel method, i.e., Adversarial Learning of Answer-Related Representation (ALARR) for visual question answering, which seeks an effective answer-related representation for the question-image pair based on adversarial learning between two processes. The embedding learning process aims to generate modality-invariant joint representations for the question-image and question-answer pairs, respectively. Meanwhile, it tries to confuse the other process, embedding discriminator, which tries to discriminate the two representations from different modalities of pairs. Specifically, the joint embedding of the question-image pair is learned by a three-level attention model, and the joint representation of the question-answer pair is learned by a semantic integration model. Through the adversarial leaning, the answer-related representation are better preserved. Then an answer predictor is proposed to infer the answer from the answer-related representation. Experiments conducted on two widely used VQA benchmark datasets demonstrate that the proposed model outperforms the state-of-the-art approaches.
Yun Liu 0017, Xiaoming Zhang 0001, Feiran Huang, Zhoujun Li 0001
CIKM1