VLDB 2026 Research / reviewers in the wild / expert
Xin Wang 0119
dblp:10/5630-119
· DBLP profile ↗
13ranked-venue papers
2as first author
10since 2021 · last 2026
0000-0001-9531-6662ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Software engineering, systems software and programming languages · 4 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Probing the Safety Robustness of LLMs in Latent SpaceabstractTianle Gu, Kexin Huang, Zongqi Wang, Yixu Wang, Jie Li, Xin Wang, Yang Yao, Yujiu Yang, Yan Teng, Yingchun Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Tianle Gu, Zongqi Wang, Yixu Wang, Jie Li 0052, Xin Wang 0119, Yujiu Yang 0001, Yan Teng 0002, Yingchun Wang 0004 |
ACL (1) | 6 |
| 2026 | NAP-Tuning: Neural Augmented Prompt Tuning for Adversarially Robust Vision-Language ModelsabstractVision-Language Models (VLMs) such as CLIP have demonstrated remarkable capabilities in understanding relationships between visual and textual data through joint embedding spaces. Despite their effectiveness, these models remain vulnerable to adversarial attacks, particularly in the image modality, posing significant security concerns. Building upon our previous work on Adversarial Prompt Tuning (AdvPT), which introduced learnable text prompts to enhance adversarial robustness in VLMs without extensive parameter training, we present a significant extension by introducing the Neural Augmentor framework for Multi-modal Adversarial Prompt Tuning (NAP-Tuning). As a significant extension, NAP-Tuning first establishes a comprehensive multi-modal (text and visual) and multi-layer prompting framework. The core of this framework is a targeted structural augmentation for feature-level purification, implemented through our Neural Augmentor approach. This framework implements feature purification by incorporating TokenRefiners-lightweight neural modules that learn to reconstruct purified features via residual connections-to directly address distortions in the feature space. This structural intervention is what enables the multi-modal and multi-layer system to effectively perform modality-specific and layer-specific feature rectification. Comprehensive experiments demonstrate that NAP-Tuning significantly outperforms existing methods across various datasets and attack types. Notably, our approach shows significant improvements over the strongest baselines under the challenging AutoAttack benchmark, outperforming them by 32.3% on ViT-B16 and 31.3% on ViT-B32 architectures while maintaining competitive clean accuracy. This work highlights the efficacy of internal feature-level intervention in prompt tuning for adversarial robustness, moving beyond input-side alignment approaches to create an adaptive defense mechanism that can identify and rectify adversarial perturbations across embedding spaces. Jiaming Zhang 0006, Xin Wang 0119, Xingjun Ma, Lingyu Qiu, Yu-Gang Jiang 0001, Jitao Sang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | TAPT: Test-Time Adversarial Prompt Tuning for Robust Inference in Vision-Language ModelsabstractLarge pre-trained Vision-Language Models (VLMs) such as CLIP have demonstrated excellent zero-shot generalizability across various downstream tasks. However, recent studies have shown that the inference performance of CLIP can be greatly degraded by small adversarial perturbations, especially its visual modality, posing significant safety threats. To mitigate this vulnerability, in this paper, we propose a novel defense method called Test-Time Adversarial Prompt Tuning (TAPT) to enhance the inference robustness of CLIP against visual adversarial attacks. TAPT is a test-time defense method that learns defensive bimodal (textual and visual) prompts to robustify the inference process of CLIP. Specifically, it is an unsupervised method that optimizes the defensive prompts for each test sample by minimizing a multi-view entropy and aligning adversarial-clean distributions. We evaluate the effectiveness of TAPT on 11 benchmark datasets, including ImageNet and 10 other zero-shot datasets, demonstrating that it enhances the zero-shot adversarial robustness of the original CLIP by at least 48.9% against AutoAttack (AA), while largely maintaining performance on clean examples. Moreover, TAPT outperforms existing adversarial prompt tuning methods across various backbones, achieving an average robustness improvement of at least 36.6%. Code is available at https://github.com/xinwong/TAPT. Xin Wang 0119, Kai Chen 0027, Jiaming Zhang 0006, Jingjing Chen 0001, Xingjun Ma |
CVPR | 1 |
| 2025 | Argus Inspection: Do Multimodal Large Language Models Possess the Eye of Panoptes?
Chiyu Chen, Zhenqi He, Yixu Wang, Xin Wang 0119, Tianle Gu, Jie Li 0052, Yan Teng 0002, Yingchun Wang 0004 |
ACM Multimedia | 7 |
| 2025 | SafeVid: Toward Safety Aligned Video Large Multimodal ModelsabstractAs Video Large Multimodal Models (VLMMs) rapidly advance, their inherent complexity introduces significant safety challenges, particularly the issue of mismatched generalization where static safety alignments fail to transfer to dynamic video contexts. We introduce SafeVid, a framework designed to instill video-specific safety principles in VLMMs. SafeVid uniquely transfers robust textual safety alignment capabilities to the video domain by employing detailed textual video descriptions as an interpretive bridge, facilitating LLM-based rule-driven safety reasoning. This is achieved through a closed-loop system comprising: 1) generation of SafeVid-350K, a novel 350,000-pair video-specific safety preference dataset; 2) targeted alignment of VLMMs using Direct Preference Optimization (DPO); and 3) comprehensive evaluation via our new SafeVidBench benchmark. Alignment with SafeVid-350K significantly enhances VLMM safety, with models like LLaVA-NeXT-Video demonstrating substantial improvements (e.g., up to 42.39%) on SafeVidBench. SafeVid provides critical resources and a structured approach, demonstrating that leveraging textual descriptions as a conduit for safety reasoning markedly improves the safety alignment of VLMMs in complex multimodal scenarios. Yixu Wang, Yifeng Gao 0002, Xin Wang 0119, Yan Teng 0002, Xingjun Ma, Yingchun Wang 0004, Yu-Gang Jiang 0001 |
NeurIPS | 4 |
| 2024 | Adversarial Prompt Tuning for Vision-Language Models
Jiaming Zhang 0006, Xingjun Ma, Xin Wang 0119, Lingyu Qiu, Jiaqi Wang 0003, Yu-Gang Jiang 0001, Jitao Sang 0001 |
ECCV (45) | 3 |
| 2024 | AdvQDet: Detecting Query-Based Adversarial Attacks with Adversarial Contrastive Prompt Tuning
Xin Wang 0119, Kai Chen 0027, Xingjun Ma, Zhineng Chen, Jingjing Chen 0001, Yu-Gang Jiang 0001 |
ACM Multimedia | 1 |
| 2024 | Navigation as Attackers Wish? Towards Building Robust Embodied Agents under Federated LearningabstractYunchao Zhang, Zonglin Di, Kaiwen Zhou, Cihang Xie, Xin Wang. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Yunchao Zhang, Zonglin Di, Kaiwen Zhou 0002, Cihang Xie, Xin Wang 0119 |
NAACL-HLT | 5 |
| 2022 | "More Than Deep Learning": post-processing for API sequence recommendation
Xin Peng 0001, Bihuan Chen 0001, Jun Sun 0001, Zhenchang Xing, Xin Wang 0119, Wenyun Zhao |
Empir. Softw. Eng. | 6 |
| 2022 | Holistic Combination of Structural and Textual Code Information for Context Based API RecommendationabstractContext based API recommendation is an important way to help developers find the needed APIs effectively and efficiently. For effective API recommendation, we need not only a joint view of both structural and textual code information, but also a holistic view of correlated API usage in control and data flow graph as a whole. Unfortunately, existing API recommendation methods exploit structural or textual code information separately. In this work, we propose a novel API recommendation approach called APIRec-CST (API Recommendation by Combining Structural and Textual code information). APIRec-CST is a deep learning model that combines the API usage with the text information in the source code based on an API Context Graph Network and a Code Token Network that simultaneously learn structural and textual features for API recommendation. We apply APIRec-CST to train a model for JDK library based on 1,914 open-source Java projects and evaluate the accuracy and MRR (Mean Reciprocal Rank) of API recommendation with another 6 open-source projects. The results show that our approach achieves respectively a top-1, top-5, top-10 accuracy and MRR of 60.3, 81.5, 87.7 and 69.4 percent, and significantly outperforms an existing graph-based statistical approach and a tree-based deep learning approach for API recommendation. A further analysis shows that textual code information makes sense and improves the accuracy and MRR. The sensitivity analysis shows that the top-k accuracy and MRR of APIRec-CST are insensitive to the number of APIs to be recommended in a hole. We also conduct a user study in which two groups of students are asked to finish 6 programming tasks with or without our APIRec-CST plugin. The results show that APIRec-CST can help the students to finish the tasks faster and more accurately and the feedback on the usability is overwhelmingly positive. Xin Peng 0001, Zhenchang Xing, Jun Sun 0001, Xin Wang 0119, Yifan Zhao 0008, Wenyun Zhao |
IEEE Trans. Software Eng. | 5 |
| 2020 | Source Code based On-demand Class Documentation GenerationabstractIn this paper, we present OpenAPIDocGen2, a tool that generates on-demand class documentation based on source code and documentation analysis. For a given class, OpenAPIDocGen2 generates a combined documentation for it, which includes functionality descriptions, directives, domain concepts, usage examples, class/method roles, key methods, relevant classes/methods, characteristics and concepts classification, and usage scenarios. Mingwei Liu 0002, Xin Peng 0001, Xiujie Meng, Huanjun Xu, Shuangshuang Xing, Xin Wang 0119, Yang Liu 0003 |
ICSME | 6 |
| 2020 | Learning based and Context Aware Non-Informative Comment DetectionabstractThis report introduces the approach that we have designed and implemented for the DeClutter challenge of Doc-Gen2, which detects non-informative code comments. The approach combines both comment based text classification and code context based prediction. Based on the approach, our "fduse" team achieved the best F1 score (0.847) in the competition. Mingwei Liu 0002, Xin Peng 0001, Chong Wang 0013, Chengyuan Zhao, Xin Wang 0119, Shuangshuang Xing |
ICSME | 6 |
| 2019 | Generative API usage code recommendation with parameter concretization
Xin Peng 0001, Jun Sun 0001, Zhenchang Xing, Xin Wang 0119, Yifan Zhao 0008, Hairui Zhang, Wenyun Zhao |
Sci. China Inf. Sci. | 5 |