VLDB 2026 Research / reviewers in the wild / expert
Youze Wang
dblp:267/2608
· DBLP profile ↗
8ranked-venue papers
5as first author
7since 2021 · last 2026
0009-0003-5621-6310ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 5 first-author · 7 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Benchmarking Trustworthiness in Multimodal LLMs for Video UnderstandingabstractRecent advancements in multimodal large language models for video understanding (videoLLMs) have enhanced their capacity to process complex spatiotemporal data. However, challenges such as factual inaccuracies, harmful content, biases, hallucinations, and privacy risks compromise their reliability. This study introduces Trust-videoLLMs, a first comprehensive benchmark evaluating 23 state-of-the-art videoLLMs (5 commercial, 18 open-source) across five critical dimensions: truthfulness, robustness, safety, fairness, and privacy. Comprising 30 tasks with adapted, synthetic, and annotated videos, the framework assesses spatiotemporal risks, temporal consistency and cross-modal impact. Results reveal significant limitations in dynamic scene comprehension, cross-modal perturbation resilience and real-world risk mitigation. While open-source models occasionally outperform, proprietary models generally exhibit superior credibility, though scaling does not consistently improve performance. These findings underscore the need for enhanced training datat diversity and robust multimodal alignment. Trust-videoLLMs provides a publicly available, extensible toolkit for standardized trustworthiness assessments, addressing the critical gap between accuracy-focused benchmarks and demands for robustness, safety, fairness, and privacy. Youze Wang, Zijun Chen 0001, Shishen Gu, Wenbo Hu 0001, Yinpeng Dong, Hang Su 0006, Jun Zhu 0001, Meng Wang 0001, Richang Hong |
AAAI | 1 |
| 2025 | Joint Adversarial Purification: Mitigating the Threat of Multimodal Adversarial ExamplesabstractVision-language pre-training (VLP) models exhibit exceptional generalization capabilities across diverse vision-language (V+L) tasks. However, studies reveal their vulnerability to carefully crafted multimodal adversarial examples. Notably, state-of-the-art attacks like Co-Attack (white-box) and SGA (transfer-based) demonstrate alarming success rates, posing critical security threats to VLP models. Current defense mechanisms against such multimodal attacks remain insufficiently explored. To address this challenge, we propose Joint Adversarial Purification (JAP), a novel defense framework that synergistically eliminates adversarial perturbations across modalities through cross-modal interaction. Our approach harnesses cross-modal semantic synergy to jointly purify adversarial perturbations: Generative denoising establishes visual-semantic anchors through diffusion processes, while purified linguistic cues conversely enhance visual perturbation filtering, forming a self-reinforcing defense cycle. Extensive experiments demonstrate that JAP effectively mitigates adversarial threats from both white-box Co-Attack and transfer-based SGA, significantly outperforming existing unimodal defense baselines. This work establishes a new paradigm for securing VLP models against multimodal adversarial attacks. Youze Wang, Wenbo Hu 0001, Richang Hong |
ICMR | 2 |
| 2025 | Making Strides Security in Multimodal Fake News Detection Models: A Comprehensive Analysis of Adversarial Attacks
Jiahua Si, Youze Wang, Wenbo Hu 0001, Qiang Liu 0006, Richang Hong |
MMM (2) | 2 |
| 2025 | Align Is Not Enough: Multimodal Universal Jailbreak Attack Against Multimodal Large Language ModelsabstractLarge Language Models (LLMs) have evolved into Multimodal Large Language Models (MLLMs), significantly enhancing their capabilities by integrating visual information and other types, thus aligning more closely with the nature of human intelligence, which processes a variety of data forms beyond just text. Despite advancements, the undesirable generation of these models remains a critical concern, particularly due to vulnerabilities exposed by text-based jailbreak attacks, which have represented a significant threat by challenging existing safety protocols. Motivated by the unique security risks posed by the integration of new and old modalities for MLLMs, we propose a unified multimodal universal jailbreak attack framework that leverages iterative image-text interactions and transfer-based strategy to generate a universal adversarial suffix and image. Our work not only highlights the interaction of image-text modalities can be used as a critical vulnerability but also validates that multimodal universal jailbreak attacks can bring higher-quality undesirable generations across different MLLMs. We evaluate the undesirable context generation of MLLMs like LLaVA, Yi-VL, MiniGPT4, MiniGPT-v2, and InstructBLIP, and reveal significant multimodal safety alignment issues, highlighting the inadequacy of current safety mechanisms against sophisticated multimodal attacks. This study underscores the urgent need for robust safety measures in MLLMs, advocating for a comprehensive review and enhancement of security protocols to mitigate potential risks associated with multimodal capabilities. Youze Wang, Wenbo Hu 0001, Yinpeng Dong, Jing Liu 0001, Hanwang Zhang, Richang Hong |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | Exploring Transferability of Multimodal Adversarial Samples for Vision-Language Pre-Training Models With Contrastive LearningabstractThe integration of visual and textual data in Vision-Language Pre-training (VLP) models is crucial for enhancing vision-language understanding. However, the adversarial robustness of these models, especially in the alignment of image-text features, has not yet been sufficiently explored. In this paper, we introduce a novel gradient-based multimodal adversarial attack method, underpinned by contrastive learning, to improve the transferability of multimodal adversarial samples in VLP models. This method concurrently generates adversarial texts and images within imperceptive perturbation, employing both image-text and intra-modal contrastive loss. We evaluate the effectiveness of our approach on image-text retrieval and visual entailment tasks, using publicly available datasets in a black-box setting. Extensive experiments indicate a significant advancement over existing single-modal transfer-based adversarial attack methods and current multimodal adversarial attack approaches. Youze Wang, Wenbo Hu 0001, Yinpeng Dong, Hanwang Zhang, Hang Su 0006, Richang Hong |
IEEE Trans. Multim. | 1 |
| 2024 | Iterative Adversarial Attack on Image-Guided Story Ending GenerationabstractMultimodal learning involves developing models that can integrate information from various sources like images and texts. In this field, multimodal text generation is a crucial aspect that involves processing data from multiple modalities and outputting text. The image-guided story ending generation (IgSEG) is a particularly significant task, targeting on an understanding of complex relationships between text and image data with a complete story text ending. Unfortunately, deep neural networks, which are the backbone of recent IgSEG models, are vulnerable to adversarial samples. Current adversarial attack methods mainly focus on single-modality data and do not analyze adversarial attacks for multimodal text generation tasks that use cross-modal information. To this end, we propose an iterative adversarial attack method (Iterative-attack) that fuses image and text modality attacks, allowing for an attack search for adversarial text and image in a more effective iterative way. Experimental results demonstrate that the proposed method outperforms existing single-modal and non-iterative multimodal attack methods, indicating the potential for improving the adversarial robustness of multimodal text generation models, such as multimodal machine translation, multimodal question answering, etc. Youze Wang, Wenbo Hu 0001, Richang Hong |
IEEE Trans. Multim. | 1 |
| 2021 | Efficient Graph Deep Learning in TensorFlow with tf_geometricabstractWe introduce tf_geometric1, an efficient and friendly library for graph deep learning, which is compatible with both TensorFlow 1.x and 2.x. It provides kernel libraries for building Graph Neural Networks (GNNs) as well as implementations of popular GNNs. The kernel libraries consist of infrastructures for building efficient GNNs, including graph data structures, graph map-reduce framework, graph mini-batch strategy, etc. These infrastructures enable tf_geometric to support single-graph computation, multi-graph computation, graph mini-batch, distributed training, etc.; therefore, tf_geometric can be used for a variety of graph deep learning tasks, such as node classification, link prediction, and graph classification. Based on the kernel libraries, tf_geometric implements a variety of popular GNN models. To facilitate the implementation of GNNs, tf_geometric also provides some other libraries for dataset management, graph sampling, etc. Different from existing popular GNN libraries, tf_geometric provides not only Object-Oriented Programming (OOP) APIs, but also Functional APIs, which enable tf_geometric to handle advanced tasks such as graph meta-learning. The APIs are friendly and suitable for both beginners and experts. Jun Hu 0016, Shengsheng Qian, Quan Fang, Youze Wang, Huaiwen Zhang, Changsheng Xu |
ACM Multimedia | 4 |
| 2020 | Fake News Detection via Knowledge-driven Multimodal Graph Convolutional NetworksabstractNowadays, with the rapid development of social media, there is a great deal of news produced every day. How to detect fake news automatically from a large of multimedia posts has become very important for people, the government and news recommendation sites. However, most of the existing approaches either extract features from the text of the post which is a single modality or simply concatenate the visual features and textual features of a post to get a multimodal feature and detect fake news. Most of them ignore the background knowledge hidden in the text content of the post which facilitates fake news detection. To address these issues, we propose a novel Knowledge-driven Multimodal Graph Convolutional Network (KMGCN) to model the semantic representations by jointly modeling the textual information, knowledge concepts and visual information into a unified framework for fake news detection. Instead of viewing text content as word sequences normally, we convert them into a graph, which can model non-consecutive phrases for better obtaining the composition of semantics. Besides, we not only convert visual information as nodes of graphs but also retrieve external knowledge from real-world knowledge graph as nodes of graphs to provide complementary semantics information to improve fake news detection. We utilize a well-designed graph convolutional network to extract the semantic representation of these graphs. Extensive experiments on two public real-world datasets illustrate the validation of our approach. Youze Wang, Shengsheng Qian, Jun Hu 0016, Quan Fang, Changsheng Xu |
ICMR | 1 |