Qingyang Teng

dblp:304/4996 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
3since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Trustworthy machine learning · 65% Vision and language · 24% Speech recognition and synthesis · 10%
Network and information security
2 papers
Security and privacy of machine learning · 100%

Topics — the 10 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Security and privacy of machine learning
adversarial attack
1.422025
Towards Visualization-of-Thought Jailbreak Attack against Large Visual Language Models · NeurIPS 2025
Black-box Adversarial Attacks on Commercial Speech Platforms with Minimal Information · CCS 2021
Computer vision › Vision and language › vision-language model
multimodal large language model
1.222026
Towards Visualization-of-Thought Jailbreak Attack against Large Visual Language Models · NeurIPS 2025
USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models · ACL (1) 2026
Machine learning › Trustworthy machine learning › generative model safety
multimodal large language model safety
1.012026
USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models · ACL (1) 2026
Machine learning › Trustworthy machine learning
safety evaluation
1.012026
USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models · ACL (1) 2026
Machine learning › Trustworthy machine learning › AI safety
safety alignment
0.912025
Towards Visualization-of-Thought Jailbreak Attack against Large Visual Language Models · NeurIPS 2025
Security and privacy of machine learning › adversarial attack
jailbreak attack
0.912025
Towards Visualization-of-Thought Jailbreak Attack against Large Visual Language Models · NeurIPS 2025
Natural language and speech › Speech recognition and synthesis
automatic speech recognition
0.512021
Black-box Adversarial Attacks on Commercial Speech Platforms with Minimal Information · CCS 2021
Security and privacy of machine learning › adversarial attack
black-box attack
0.512021
Black-box Adversarial Attacks on Commercial Speech Platforms with Minimal Information · CCS 2021
Security and privacy of machine learning › adversarial attack
physical adversarial attack
0.512021
Black-box Adversarial Attacks on Commercial Speech Platforms with Minimal Information · CCS 2021
Machine learning › Trustworthy machine learning
robustness
0.312025
Towards Visualization-of-Thought Jailbreak Attack against Large Visual Language Models · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

visualization-of-thought · 1.7chain-of-thought · 1.7safety benchmarking · 1.0model inversion · 1.0decision-only optimization · 1.0adaptive problem decomposition · 1.0
YearPublicationVenuePosition
2026 USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models
abstract
Baolin Zheng, Guanlin Chen, Qingyang Teng, Hongqiong Zhong, Yingshui Tan, Zhendong Liu, Weixun Wang, Jiaheng Liu, Jian Yang, Huiyun Jing, Jincheng Wei, Wenbo Su, Xiaoyong Zhu, Bo Zheng, Kaifu Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Baolin Zheng, Qingyang Teng, Hongqiong Zhong, Yingshui Tan, Weixun Wang, Jian Yang 0037, Huiyun Jing, Jincheng Wei, Wenbo Su, Xiaoyong Zhu, Bo Zheng 0007, Kaifu Zhang
ACL (1)3
2025 Towards Visualization-of-Thought Jailbreak Attack against Large Visual Language Models
abstract
As Visual Language Models (VLMs) continue to evolve, they have demonstrated increasingly sophisticated logical reasoning capabilities and multimodal thought generation, opening doors to widespread applications. However, this advancement raises serious concerns about content security, particularly when these models process complex multimodal inputs requiring intricate reasoning. When faced with these safety challenges, the critical competition between logical reasoning and safety objectives of VLMs is often overlooked in previous works. In this paper, we introduce Visualization-of-Thought Attack (\textbf{VoTA}), a novel and automated attack framework that strategically constructs chains of images with risky visual thoughts to challenge victim models. Our attack provokes the inherent conflict between the model's logical processing and safety protocols, ultimately leading to the generation of unsafe content. Through comprehensive experiments, VoTA achieves remarkable effectiveness, improving the average attack success rate (ASR) by 26.71\% (from 63.70\% to 90.41\%) on 9 open-source and 6 commercial VLMs, compared to the state-of-the-art methods. These results expose a critical vulnerability: current VLMs struggle to maintain safety guarantees when processing insecure multimodal visualization-of-thought inputs, highlighting the urgency and necessity of enhancing safety alignment. Our code and dataset are available at https://github.com/Hongqiong12/VoTA. Content Warning: This paper contains harmful contents that may be offensive.
Hongqiong Zhong, Qingyang Teng, Baolin Zheng, Yingshui Tan, Wenbo Su, Xiaoyong Zhu, Bo Zheng 0007, Kaifu Zhang
NeurIPS2
2021 Black-box Adversarial Attacks on Commercial Speech Platforms with Minimal Information
abstract
Adversarial attacks against commercial black-box speech platforms, including cloud speech APIs and voice control devices, have received little attention until recent years. Constructing such attacks is difficult mainly due to the unique characteristics of time-domain speech signals and the much more complex architecture of acoustic systems. The current "black-box" attacks all heavily rely on the knowledge of prediction/confidence scores or other probability information to craft effective adversarial examples (AEs), which can be intuitively defended by service providers without returning these messages. In this paper, we take one more step forward and propose two novel adversarial attacks in more practical and rigorous scenarios. For commercial cloud speech APIs, we propose Occam, a decision-only black-box adversarial attack, where only final decisions are available to the adversary. In Occam, we formulate the decision-only AE generation as a discontinuous large-scale global optimization problem, and solve it by adaptively decomposing this complicated problem into a set of sub-problems and cooperatively optimizing each one. Our Occam is a one-size-fits-all approach, which achieves 100% success rates of attacks (SRoA) with an average SNR of 14.23dB, on a wide range of popular speech and speaker recognition APIs, including Google, Alibaba, Microsoft, Tencent, iFlytek, and Jingdong, outperforming the state-of-the-art black-box attacks. For commercial voice control devices, we propose NI-Occam, the first non-interactive physical adversarial attack, where the adversary does not need to query the oracle and has no access to its internal information and training data. We, for the first time, combine adversarial attacks with model inversion attacks, and thus generate the physically-effective audio AEs with high transferability without any interaction with target devices. Our experimental results show that NI-Occam can successfully fool Apple Siri, Microsoft Cortana, Google Assistant, iFlytek and Amazon Echo with an average SRoA of 52% and SNR of 9.65dB, shedding light on non-interactive physical attacks against voice control devices.
Baolin Zheng, Peipei Jiang 0002, Qian Wang 0002, Qi Li 0002, Chao Shen 0001, Cong Wang 0001, Yunjie Ge, Qingyang Teng, Shenyi Zhang
CCS8