Yingshui Tan

dblp:245/8953 · DBLP profile ↗
← Back
10ranked-venue papers
2as first author
9since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 2 first-author · 9 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Trustworthy machine learning · 42% Language models and text generation · 40% Vision and language · 9%
Network and information security
2 papers
Security and privacy of machine learning · 100%

Topics — the 16 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
safety evaluation
1.922026
USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models · ACL (1) 2026
Chinese SafetyQA: A Safety Short-form Factuality Benchmark for Large Language Models · ACL (1) 2025
Security and privacy of machine learning
adversarial attack
1.722025
Towards Visualization-of-Thought Jailbreak Attack against Large Visual Language Models · NeurIPS 2025
HiddenDetect: Detecting Jailbreak Attacks against Multimodal Large Language Models via Monitoring Hidden States · ACL (1) 2025
Security and privacy of machine learning › adversarial attack
jailbreak attack
1.722025
Towards Visualization-of-Thought Jailbreak Attack against Large Visual Language Models · NeurIPS 2025
HiddenDetect: Detecting Jailbreak Attacks against Multimodal Large Language Models via Monitoring Hidden States · ACL (1) 2025
Computer vision › Vision and language › vision-language model
multimodal large language model
1.222026
Towards Visualization-of-Thought Jailbreak Attack against Large Visual Language Models · NeurIPS 2025
USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models · ACL (1) 2026
Machine learning › Trustworthy machine learning
robustness
1.122025
HiddenDetect: Detecting Jailbreak Attacks against Multimodal Large Language Models via Monitoring Hidden States · ACL (1) 2025
Towards Visualization-of-Thought Jailbreak Attack against Large Visual Language Models · NeurIPS 2025
Machine learning › Trustworthy machine learning › generative model safety
multimodal large language model safety
1.012026
USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models · ACL (1) 2026
Natural language and speech › Language models and text generation
code generation
0.912025
M2RC-EVAL: Massively Multilingual Repository-level Code Completion Evaluation · ACL (1) 2025
Natural language and speech › Language models and text generation › evaluation of language models
factuality evaluation
0.912025
Chinese SimpleQA: A Chinese Factuality Evaluation for Large Language Models · ACL (1) 2025
Natural language and speech › Language models and text generation › large language model evaluation › capability evaluation
game-based evaluation
0.912025
KORGym: A Dynamic Game Platform for LLM Reasoning Evaluation · NeurIPS 2025
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
game playing
0.912025
KORGym: A Dynamic Game Platform for LLM Reasoning Evaluation · NeurIPS 2025
Machine learning › Trustworthy machine learning › AI safety
jailbreak detection
0.912025
HiddenDetect: Detecting Jailbreak Attacks against Multimodal Large Language Models via Monitoring Hidden States · ACL (1) 2025
Natural language and speech › Language models and text generation
large language model
0.912025
KORGym: A Dynamic Game Platform for LLM Reasoning Evaluation · NeurIPS 2025
Natural language and speech › Language models and text generation
multilingual language models
0.912025
M2RC-EVAL: Massively Multilingual Repository-level Code Completion Evaluation · ACL (1) 2025
Natural language and speech › Language models and text generation › evaluation of language models
reasoning evaluation
0.912025
KORGym: A Dynamic Game Platform for LLM Reasoning Evaluation · NeurIPS 2025
Machine learning › Trustworthy machine learning › AI safety
safety alignment
0.912025
Towards Visualization-of-Thought Jailbreak Attack against Large Visual Language Models · NeurIPS 2025
Natural language and speech › Language models and text generation
large language model evaluation
0.312025
Chinese SafetyQA: A Safety Short-form Factuality Benchmark for Large Language Models · ACL (1) 2025

Methods — techniques the papers use, named apart from their topics

visualization-of-thought · 1.7hidden state monitoring · 1.7chain-of-thought · 1.7activation analysis · 1.7benchmark construction · 1.7safety benchmarking · 1.0reinforcement learning scenarios · 0.9interactive evaluation · 0.9
YearPublicationVenuePosition
2026 USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models
abstract
Baolin Zheng, Guanlin Chen, Qingyang Teng, Hongqiong Zhong, Yingshui Tan, Zhendong Liu, Weixun Wang, Jiaheng Liu, Jian Yang, Huiyun Jing, Jincheng Wei, Wenbo Su, Xiaoyong Zhu, Bo Zheng, Kaifu Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Baolin Zheng, Qingyang Teng, Hongqiong Zhong, Yingshui Tan, Weixun Wang, Jian Yang 0037, Huiyun Jing, Jincheng Wei, Wenbo Su, Xiaoyong Zhu, Bo Zheng 0007, Kaifu Zhang
ACL (1)5
2025 Chinese SimpleQA: A Chinese Factuality Evaluation for Large Language Models
abstract
Yancheng He, Shilong Li, Jiaheng Liu, Yingshui Tan, Weixun Wang, Hui Huang, Xingyuan Bu, Hangyu Guo, Chengwei Hu, Boren Zheng, Zhuoran Lin, Dekai Sun, Zhicheng Zheng, Wenbo Su, Bo Zheng. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Yancheng He, Yingshui Tan, Weixun Wang, Hui Huang 0021, Xingyuan Bu, Hangyu Guo, Chengwei Hu, Boren Zheng, Zhuoran Lin, Dekai Sun, Zhicheng Zheng, Wenbo Su, Bo Zheng 0007
ACL (1)4
2025 HiddenDetect: Detecting Jailbreak Attacks against Multimodal Large Language Models via Monitoring Hidden States
abstract
The integration of additional modalities increases the susceptibility of large vision-language models (LVLMs) to safety risks, such as jailbreak attacks, compared to their language-only counterparts. While existing research primarily focuses on post-hoc alignment techniques, the underlying safety mechanisms within LVLMs remain largely unexplored. In this work , we investigate whether LVLMs inherently encode safety-relevant signals within their internal activations during inference. Our findings reveal that LVLMs exhibit distinct activation patterns when processing unsafe prompts, which can be leveraged to detect and mitigate adversarial inputs without requiring extensive fine-tuning. Building on this insight, we introduce HiddenDetect, a novel tuning-free framework that harnesses internal model activations to enhance safety. Experimental results show that HiddenDetect surpasses state-of-the-art methods in detecting jailbreak attacks against LVLMs. By utilizing intrinsic safety-aware patterns, our method provides an efficient and scalable solution for strengthening LVLM robustness against multimodal threats. Our code and data will be released publicly.
Yilei Jiang, Xinyan Gao, Tianshuo Peng, Yingshui Tan, Xiaoyong Zhu, Bo Zheng 0007, Xiangyu Yue 0001
ACL (1)4
2025 M2RC-EVAL: Massively Multilingual Repository-level Code Completion Evaluation
abstract
Jiaheng Liu, Ken Deng, Congnan Liu, Jian Yang, Shukai Liu, He Zhu, Peng Zhao, Linzheng Chai, Yanan Wu, JinKe JinKe, Ge Zhang, Zekun Moore Wang, Guoan Zhang, Yingshui Tan, Bangyu Xiang, Zhaoxiang Zhang, Wenbo Su, Bo Zheng. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Ken Deng, Congnan Liu, Jian Yang 0030, Linzheng Chai, Ge Zhang 0009, Zekun Moore Wang, Guoan Zhang, Yingshui Tan, Bangyu Xiang, Zhaoxiang Zhang 0001, Wenbo Su, Bo Zheng 0007
ACL (1)14
2025 Chinese SafetyQA: A Safety Short-form Factuality Benchmark for Large Language Models
abstract
Yingshui Tan, Boren Zheng, Baihui Zheng, Kerui Cao, Huiyun Jing, Jincheng Wei, Jiaheng Liu, Yancheng He, Wenbo Su, Xiaoyong Zhu, Bo Zheng, Kaifu Zhang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Yingshui Tan, Boren Zheng, Baihui Zheng, Kerui Cao, Huiyun Jing, Jincheng Wei, Yancheng He, Wenbo Su, Xiaoyong Zhu, Bo Zheng 0007, Kaifu Zhang
ACL (1)1
2025 Multi-Step Adaptive Attack Agent: A Dynamic Approach for Jailbreaking Large Language Models
abstract
Large Language Models (LLMs) have showcased remarkable potential across various domains, especially in text generation. However, their vulnerability to jailbreak attacks presents considerable challenges to secure deployment, as attackers can use carefully crafted prompts to bypass safety measures and generate harmful content. Current jailbreak methods generally suffer from two significant limitations: a restricted strategy space for generating adversarial prompts and insufficient optimization of prompts based on feedback from LLMs. To overcome these challenges, we present Multistep Adaptive Attack Agent (MATA), an approach that employs a game-theoretic interaction between attack model and target model to adaptively execute jailbreak attacks on LLMs. This method enables iterative attempts based on reflection, gradually identifying the optimal jailbreak attack strategy within a complex strategy space. We compared MATA with mainstream methods across multiple open-source and closed-source LLMs, including Llama3.1, GLM4, and GPT4o. The results demonstrate that our approach exceeds existing methods in terms of attack success rate, average number of queries, and prompt diversity, effectively identifying vulnerabilities in LLMs.
Huiyun Jing, Jincheng Wei, Yingshui Tan, Boren Zheng, Qingsong Yao
ICTAI4
2025 KORGym: A Dynamic Game Platform for LLM Reasoning Evaluation
abstract
Recent advancements in large language models (LLMs) underscore the need for more comprehensive evaluation methods to accurately assess their reasoning capabilities. Existing benchmarks are often domain-specific and thus cannot fully capture an LLM’s general reasoning potential. To address this limitation, we introduce the **Knowledge Orthogonal Reasoning Gymnasium (KORGym)**, a dynamic evaluation platform inspired by KOR-Bench and Gymnasium. KORGym offers over fifty games in either textual or visual formats and supports interactive, multi-turn assessments with reinforcement learning scenarios. Using KORGym, we conduct extensive experiments on 19 LLMs and 8 VLMs, revealing consistent reasoning patterns within model families and demonstrating the superior performance of closed-source models. Further analysis examines the effects of modality, reasoning strategies, reinforcement learning techniques, and response length on model performance. We expect KORGym to become a valuable resource for advancing LLM reasoning research and developing evaluation methodologies suited to complex, interactive environments.
Jiajun Shi, Jian Yang 0037, Xingyuan Bu, Jiangjie Chen, Junting Zhou, Kaijing Ma, Zhoufutu Wen, Bingli Wang, Yancheng He, Hualei Zhu, Wei Zhang 0021, Ruibin Yuan, Yunli Wang, Siyuan Fang, Qianyu He, Robert Tang, Yingshui Tan, Wangchunshu Zhou, Zhaoxiang Zhang 0001, Zhoujun Li 0001, Wenhao Huang 0001, Ge Zhang 0009
NeurIPS24
2025 Towards Visualization-of-Thought Jailbreak Attack against Large Visual Language Models
abstract
As Visual Language Models (VLMs) continue to evolve, they have demonstrated increasingly sophisticated logical reasoning capabilities and multimodal thought generation, opening doors to widespread applications. However, this advancement raises serious concerns about content security, particularly when these models process complex multimodal inputs requiring intricate reasoning. When faced with these safety challenges, the critical competition between logical reasoning and safety objectives of VLMs is often overlooked in previous works. In this paper, we introduce Visualization-of-Thought Attack (\textbf{VoTA}), a novel and automated attack framework that strategically constructs chains of images with risky visual thoughts to challenge victim models. Our attack provokes the inherent conflict between the model's logical processing and safety protocols, ultimately leading to the generation of unsafe content. Through comprehensive experiments, VoTA achieves remarkable effectiveness, improving the average attack success rate (ASR) by 26.71\% (from 63.70\% to 90.41\%) on 9 open-source and 6 commercial VLMs, compared to the state-of-the-art methods. These results expose a critical vulnerability: current VLMs struggle to maintain safety guarantees when processing insecure multimodal visualization-of-thought inputs, highlighting the urgency and necessity of enhancing safety alignment. Our code and dataset are available at https://github.com/Hongqiong12/VoTA. Content Warning: This paper contains harmful contents that may be offensive.
Hongqiong Zhong, Qingyang Teng, Baolin Zheng, Yingshui Tan, Wenbo Su, Xiaoyong Zhu, Bo Zheng 0007, Kaifu Zhang
NeurIPS5
2021 One-class graph neural networks for anomaly detection in attributed networks
Xuhong Wang, Baihong Jin, Ping Cui, Yingshui Tan, Yupu Yang
Neural Comput. Appl.5
2019 An Encoder-Decoder Based Approach for Anomaly Detection with Application in Additive Manufacturing
abstract
We present a novel unsupervised deep learning approach that utilizes an encoder-decoder architecture for detecting anomalies in sequential sensor data collected during industrial manufacturing. Our approach is designed to not only detect whether there exists an anomaly at a given time step, but also to predict what will happen next in the (sequential) process. We demonstrate our approach on a dataset collected from a real-world Additive Manufacturing (AM) testbed. The dataset contains infrared (IR) images collected under both normal conditions and synthetic anomalies. We show that our encoder-decoder model is able to identify the injected anomalies in a modern AM manufacturing process in an unsupervised fashion. In addition, our approach also gives hints about the temperature non-uniformity of the testbed during manufacturing, which was not previously known prior to the experiment.
Yingshui Tan, Baihong Jin, Alexander J. Nettekoven, Yuxin Chen 0001, Yisong Yue, Ufuk Topcu, Alberto L. Sangiovanni-Vincentelli
ICMLA1