VLDB 2026 Research / reviewers in the wild / expert
Shenyuan Zhang
dblp:314/0011
· DBLP profile ↗
12ranked-venue papers
2as first author
12since 2021 · last 2026
0000-0002-8613-7651ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards social-aware image captioning via chain-of-thought prompting
Shenyuan Zhang, Ning Xu 0003, Quanhan Wu, Jinlin Guo, Hongshuo Tian, Anan Liu |
Expert Syst. Appl. | 1 |
| 2026 | Node Injection-Based Adversarial Attack and Defense on Social Bot DetectionabstractSocial platforms such as Twitter are increasingly threatened by automated social bots, which can manipulate public opinion, spread misinformation, and compromise platform integrity. To detect such accounts, various methods have been proposed, with many recent approaches leveraging graph neural networks (GNNs) due to their ability to model relational structures among users. However, these models may be vulnerable to adversarial perturbations that exploit their structural dependencies. To investigate this vulnerability, we propose a targeted black-box node injection attack that deceives detection models by injecting a new bot near a target bot, causing both to evade detection. This attack is specifically designed to exploit GNN-based model behaviors and does not require any access to model parameters or architecture. Given the success of this attack in exposing structural weaknesses, it is essential to develop effective defenses to improve the robustness of detection models. To this end, we introduce a cost-efficient adversarial training method with reweighted supervision, which selectively emphasizes adversarial samples during learning. This approach improves model robustness without increasing computational cost or requiring additional data. The effectiveness of our attack and defense methods is demonstrated through experiential evaluations conducted on six different models. Specifically, we achieve a maximum attack success rate (ASR) of 95.74% and robustness improvement of 95.74% in Cresci-2015 and achieve a maximum ASR of 97.15% and robustness improvement of 56.20% in TwiBot-22. The detection performance of the models fluctuates less than 2% after adversarial training. Yanwei Xie, Weizhi Nie, Lanjun Wang, Xinran Qiao, Shenyuan Zhang, Anan Liu |
IEEE Trans. Comput. Soc. Syst. | 5 |
| 2026 | How to Understand Named Entities: Using Commonsense for News CaptioningabstractNews captioning aims to describe an image with its news article body as input. It greatly relies on a set of detected named entities, including real-world people, organizations, and places. This article exploits commonsense knowledge to understand named entities for news captioning. By “understand,” we mean correlating the news content with commonsense in the wild, which helps an agent to (1) distinguish semantically similar named entities and (2) describe named entities using words outside of training corpora. Our approach consists of three modules: (a) Filter Module aims to clarify the commonsense concerning a named entity from two aspects: what does it mean ? and what is it related to ?, which divide the commonsense into explanatory knowledge and relevant knowledge , respectively. (b) Distinguish Module aggregates explanatory knowledge from node-degree , dependency , and distinguish three aspects to distinguish semantically similar named entities. (c) Enrich Module attaches relevant knowledge to named entities to enrich the entity description by commonsense information (e.g., identity and social position). Finally, all of information is integrated into the large multimodal model to generate the news caption. Extensive experiments on two challenging datasets (i.e., GoodNews and NYTimes) demonstrate the superiority of our method. Ablation studies and visualization further validate its effectiveness in understanding named entities. Shenyuan Zhang, Ning Xu 0003, Yanhui Wang 0001, Tongle Ma, Wu Liu 0005, Jinlin Guo, Anan Liu |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2024 | Beyond Users: Denoising Behavior-based Contrastive Learning for Disentangled Cross-Domain Recommendation
Lele Sun, Jing Liu 0002, Shenyuan Zhang, Weizhi Nie, Anan Liu, Yuting Su 0001 |
DASFAA (2) | 3 |
| 2024 | Prior knowledge guided text to image generation
Anan Liu, Zefang Sun, Ning Xu 0003, Rongbao Kang, Jinbo Cao, Weijun Qin, Shenyuan Zhang, Xuanya Li |
Pattern Recognit. Lett. | 8 |
| 2023 | Exploring visual relationship for social media popularity prediction
Anan Liu, Ning Xu 0003, Shenyuan Zhang, Yejun Tang, Xuanya Li |
J. Vis. Commun. Image Represent. | 5 |
| 2023 | SMPC: boosting social media popularity prediction with caption
Anan Liu, Ning Xu 0003, Jing Liu 0002, Yuting Su 0001, Shenyuan Zhang, Yejun Tang, Junbo Guo, Guoqing Jin, Xuanya Li |
Multim. Syst. | 7 |
| 2022 | LS-GAN: Iterative Language-based Image Manipulation via Long and Short Term Consistency ReasoningabstractIterative language-based image manipulation aims to edit images step by step according to user's linguistic instructions. The existing methods mostly focus on aligning the attributes and appearance of new-added visual elements with current instruction. However, they fail to maintain consistency between instructions and images as iterative rounds increase. To address this issue, we propose a novel Long and Short term consistency reasoning Generative Adversarial Network (LS-GAN), which enhances the awareness of previous objects with current instruction and better maintains the consistency with the user's intent under the continuous iterations. Specifically, we first design a Context-aware Phrase Encoder (CPE) to learn the user's intention by extracting different phrase-level information about the instruction. Further, we introduce a Long and Short term Consistency Reasoning (LSCR) mechanism. The long-term reasoning improves the model on semantic understanding and positional reasoning, while short-term reasoning ensures the ability to construct visual scenes based on linguistic instructions. Extensive results show that LS-GAN improves the generation quality in terms of both object identity and position, and achieves the state-of-the-art performance on two public datasets. Gaoxiang Cong 0001, Liang Li 0003, Zhenhuan Liu, Yunbin Tu, Weijun Qin, Shenyuan Zhang, Chengang Yan, Bin Jiang 0011 |
ACM Multimedia | 6 |
| 2022 | Improved Semantic Representation Learning by Multiple Clustering for Image-Based 3D Model RetrievalabstractUnder the heavy management on the increasing 3D models, the topic of image-based 3D model retrieval which organizes unlabeled 3D models based on abundant knowledge learned from labeled 2D images has drawn attention. However, prior methods are limited in aligning semantically at corresponding categories of two domains due to the lack of label information in the 3D domain. To this end, this paper proposes an improved semantic representation learning by multiple clustering approach, which improves the reliability of pseudo labels for 3D models, so as to achieve class-level semantic alignment. Specifically, this paper first extracts features for 2D images and 3D models. Then it clusters combining the 3D features with the semantic information from multiple clustering on 3D model features to obtain more reliable target pseudo labels. Extensive experiments have shown that the proposed method has achieved the gain of 3.0%-205.0% averagely for popular retrieval metrics on the benchmark of monocular image-based 3D object retrieval (MI3DOR), and 1.3%-69.7% on another advanced benchmark, MI3DOR-2. Jinghui Chu, Xiaoqian Zhao, Dan Song 0006, Wenhui Li 0001, Shenyuan Zhang, Xuanya Li, Anan Liu |
Int. J. Semantic Web Inf. Syst. | 5 |
| 2022 | Collaborative Distribution Alignment for 2D image-based 3D shape retrieval
Nian Hu, Heyu Zhou, Anan Liu, Xiangdong Huang 0002, Shenyuan Zhang, Guoqing Jin, Junbo Guo, Xuanya Li |
J. Vis. Commun. Image Represent. | 5 |
| 2022 | A review of feature fusion-based media popularity prediction methodsabstractWith the popularization of social media, the way of information transmission has changed, and the prediction of information popularity based on social media platforms has attracted extensive attention. Feature fusion-based media popularity prediction methods focus on the multi-modal features of social media, which aim at exploring the key factors affecting media popularity. Meanwhile, the methods make up for the deficiency in feature utilization of traditional methods based on information propagation processes. In this paper, we review feature fusion-based media popularity prediction methods from the perspective of feature extraction and predictive model construction. Before that, we analyze the influencing factors of media popularity to provide intuitive understanding. We further argue about the advantages and disadvantages of existing methods and datasets to highlight the future directions. Finally, we discuss the applications of popularity prediction. To the best of our knowledge, this is the first survey reporting feature fusion-based media popularity prediction methods. Anan Liu, Ning Xu 0003, Junbo Guo, Guoqing Jin, Yejun Tang, Shenyuan Zhang |
Vis. Informatics | 8 |
| 2021 | Image Captioning with multi-level similarity-guided semantic matchingabstractImage Captioning is a cross-modal task that needs to automatically generate coherent natural sentences to describe the image contents. Due to the large gap between vision and language modalities, most of the existing methods have the problem of inaccurate semantic matching between images and generated captions. To solve the problem, this paper proposes a novel multi-level similarity-guided semantic matching method for image captioning, which can fuse local and global semantic similarities to learn the latent semantic correlation between images and generated captions. Specifically, we extract the semantic units containing fine-grained semantic information of images and generated captions, respectively. Based on the comparison of the semantic units, we design a local semantic similarity evaluation mechanism. Meanwhile, we employ the CIDEr score to characterize the global semantic similarity. The local and global two-level similarities are finally fused using the reinforcement learning theory, to guide the model optimization to obtain better semantic matching. The quantitative and qualitative experiments on large-scale MSCOCO dataset illustrate the superiority of the proposed method, which can achieve fine-grained semantic matching of images and generated captions. Jiesi Li, Ning Xu 0003, Weizhi Nie, Shenyuan Zhang |
Vis. Informatics | 4 |