Yichi Zhang 0012

dblp:86/7054-12 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
9since 2021 · last 2025
0000-0002-1894-3977ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 4 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Trustworthy machine learning · 61% Language models and text generation · 15% Vision and language · 10%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computational science and engineering · 100%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Performance modeling and evaluation · 100%

Topics — the 28 heaviest of 29, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
robustness
2.332025
STAIR: Improving Safety Alignment with Introspective Reasoning · ICML 2025
MultiTrust: A Comprehensive Benchmark Towards Trustworthy Multimodal Large Language Models · NeurIPS 2024
Understanding the Robustness of 3D Object Detection with Bird'View Representations in Autonomous Driving · CVPR 2023
Machine learning › Trustworthy machine learning › robustness
adversarial attack
1.522024
MultiTrust: A Comprehensive Benchmark Towards Trustworthy Multimodal Large Language Models · NeurIPS 2024
Rethinking Model Ensemble in Transfer-based Adversarial Attacks · ICLR 2024
Machine learning › Trustworthy machine learning › robustness
adversarial robustness
1.422024
Rethinking Model Ensemble in Transfer-based Adversarial Attacks · ICLR 2024
Understanding the Robustness of 3D Object Detection with Bird'View Representations in Autonomous Driving · CVPR 2023
Computer vision › Vision and language › vision-language model
multimodal large language model
1.022024
Exploring the Transferability of Visual Prompting for Multimodal Large Language Models · CVPR 2024
MultiTrust: A Comprehensive Benchmark Towards Trustworthy Multimodal Large Language Models · NeurIPS 2024
Natural language and speech › Language models and text generation › large language model
emergent abilities
0.912025
DeceptionBench: A Comprehensive Benchmark for AI Deception Behaviors in Real-world Scenarios · NeurIPS 2025
Machine learning › Trustworthy machine learning
interpretability
0.912025
Mitigating Overthinking in Large Reasoning Models via Manifold Steering · NeurIPS 2025
Knowledge, reasoning and agents › Knowledge representation and reasoning › epistemic reasoning
introspective reasoning
0.912025
STAIR: Improving Safety Alignment with Introspective Reasoning · ICML 2025
Machine learning › Trustworthy machine learning › adversarial machine learning › adversarial defense
jailbreak defense
0.912025
STAIR: Improving Safety Alignment with Introspective Reasoning · ICML 2025
Natural language and speech › Language models and text generation
large language model
0.912025
DeceptionBench: A Comprehensive Benchmark for AI Deception Behaviors in Real-world Scenarios · NeurIPS 2025
Machine learning › Trustworthy machine learning › interpretability
mechanistic interpretability
0.912025
Mitigating Overthinking in Large Reasoning Models via Manifold Steering · NeurIPS 2025
Natural language and speech › Language models and text generation › large language model reasoning › efficient reasoning
overthinking mitigation
0.912025
Mitigating Overthinking in Large Reasoning Models via Manifold Steering · NeurIPS 2025
Machine learning › Trustworthy machine learning › AI safety
safety alignment
0.912025
STAIR: Improving Safety Alignment with Introspective Reasoning · ICML 2025
Machine learning › Trustworthy machine learning › robustness › adversarial robustness
adversarial transferability
0.812024
Rethinking Model Ensemble in Transfer-based Adversarial Attacks · ICLR 2024
Machine learning › Transfer learning and domain adaptation › knowledge transfer
cross-model knowledge transfer
0.812024
Exploring the Transferability of Visual Prompting for Multimodal Large Language Models · CVPR 2024
Machine learning › Trustworthy machine learning › robustness › adversarial attack
ensemble attack
0.812024
Rethinking Model Ensemble in Transfer-based Adversarial Attacks · ICLR 2024
Machine learning › Trustworthy machine learning
privacy
0.812024
MultiTrust: A Comprehensive Benchmark Towards Trustworthy Multimodal Large Language Models · NeurIPS 2024
Machine learning › Trustworthy machine learning › privacy
privacy leakage
0.812024
MultiTrust: A Comprehensive Benchmark Towards Trustworthy Multimodal Large Language Models · NeurIPS 2024
Computer vision › Vision and language
visual prompting
0.812024
Exploring the Transferability of Visual Prompting for Multimodal Large Language Models · CVPR 2024
Computational science and engineering
partial differential equation solver
0.812024
PINNacle: A Comprehensive Benchmark of Physics-Informed Neural Networks for Solving PDEs · NeurIPS 2024
Computational science and engineering › scientific machine learning › physics-informed machine learning › physics-informed neural networks
partial differential equation solving
0.812024
PINNacle: A Comprehensive Benchmark of Physics-Informed Neural Networks for Solving PDEs · NeurIPS 2024
Computational science and engineering › scientific machine learning › physics-informed machine learning
physics-informed neural networks
0.812024
PINNacle: A Comprehensive Benchmark of Physics-Informed Neural Networks for Solving PDEs · NeurIPS 2024
Computational science and engineering
scientific machine learning
0.812024
PINNacle: A Comprehensive Benchmark of Physics-Informed Neural Networks for Solving PDEs · NeurIPS 2024
Performance modeling and evaluation
benchmarking
0.812024
PINNacle: A Comprehensive Benchmark of Physics-Informed Neural Networks for Solving PDEs · NeurIPS 2024
Computer vision › 3D vision
3d object detection
0.712023
Understanding the Robustness of 3D Object Detection with Bird'View Representations in Autonomous Driving · CVPR 2023
Computer vision › 3D vision › 3d scene understanding
bird's-eye-view representation
0.712023
Understanding the Robustness of 3D Object Detection with Bird'View Representations in Autonomous Driving · CVPR 2023
Natural language and speech › Language models and text generation
chain-of-thought reasoning
0.312025
Mitigating Overthinking in Large Reasoning Models via Manifold Steering · NeurIPS 2025
Natural language and speech › Language models and text generation
preference optimization
0.312025
STAIR: Improving Safety Alignment with Introspective Reasoning · ICML 2025
Computer vision › Vision and language
multimodal reasoning
0.212024
Exploring the Transferability of Visual Prompting for Multimodal Large Language Models · CVPR 2024

Methods — techniques the papers use, named apart from their topics

benchmark construction · 1.6loss reweighting · 1.5domain decomposition · 1.5process reward model · 0.9preference optimization · 0.9multi-turn interaction evaluation · 0.9monte carlo tree search · 0.9manifold projection · 0.9chain-of-thought reasoning · 0.9activation steering · 0.9task semantics enrichment · 0.8feature consistency alignment · 0.8
YearPublicationVenuePosition
2025 STAIR: Improving Safety Alignment with Introspective Reasoning
abstract
Ensuring the safety and harmlessness of Large Language Models (LLMs) has become equally critical as their performance in applications. However, existing safety alignment methods typically suffer from safety-performance trade-offs and susceptibility to jailbreak attacks, primarily due to their reliance on direct refusals for malicious queries. In this paper, we propose STAIR, a novel framework that integrates SafeTy Alignment with Itrospective Reasoning. We enable LLMs to identify safety risks through step-by-step analysis by self-improving chain-of-thought (CoT) reasoning with safety awareness. STAIR first equips the model with a structured reasoning capability and then advances safety alignment via iterative preference optimization on step-level reasoning data generated using our newly proposed Safety-Informed Monte Carlo Tree Search (SI-MCTS). Specifically, we design a theoretically grounded reward for outcome evaluation to seek balance between helpfulness and safety. We further train a process reward model on this data to guide test-time searches for improved responses. Extensive experiments show that STAIR effectively mitigates harmful outputs while better preserving helpfulness, compared to instinctive alignment strategies. With test-time scaling, STAIR achieves a safety performance comparable to Claude-3.5 against popular jailbreak attacks. We have open-sourced our code, datasets and models at https://github.com/thu-ml/STAIR.
Yichi Zhang 0012, Zeyu Xia 0003, Zhengwei Fang, Xiao Yang 0028, Ranjie Duan, Yinpeng Dong, Jun Zhu 0001
ICML1
2025 Mitigating Overthinking in Large Reasoning Models via Manifold Steering
abstract
Recent advances in Large Reasoning Models (LRMs) have demonstrated remarkable capabilities in solving complex tasks such as mathematics and coding. However, these models frequently exhibit a phenomenon known as *overthinking* during inference, characterized by excessive validation loops and redundant deliberation, leading to substantial computational overheads. In this paper, we aim to mitigate overthinking by investigating the underlying mechanisms from the perspective of mechanistic interpretability. We first showcase that the tendency of overthinking can be effectively captured by a single direction in the model's activation space and the issue can be eased by intervening the activations along this direction. However, this efficacy soon reaches a plateau and even deteriorates as the intervention strength increases. We therefore systematically explore the activation space and find that the overthinking phenomenon is actually tied to a low-dimensional manifold, which indicates that the limited effect stems from the noises introduced by the high-dimensional steering direction. Based on this insight, we propose **Manifold Steering**, a novel approach that elegantly projects the steering direction onto the low-dimensional activation manifold given the theoretical approximation of the interference noise. Extensive experiments on DeepSeek-R1 distilled models validate that our method reduces output tokens by up to 71\% while maintaining and even improving the accuracy on several mathematical benchmarks. Our method also exhibits robust cross-domain transferability, delivering consistent token reduction performance in code generation and knowledge-based QA tasks. Code is available at: https://github.com/Aries-iai/Manifold_Steering.
Huanran Chen, Shouwei Ruan, Yichi Zhang 0012, Xingxing Wei 0001, Yinpeng Dong
NeurIPS4
2025 DeceptionBench: A Comprehensive Benchmark for AI Deception Behaviors in Real-world Scenarios
abstract
Despite the remarkable advances of Large Language Models (LLMs) across diverse cognitive tasks, the rapid enhancement of these capabilities also introduces emergent deception behaviors that may induce severe risks in high-stakes deployments. More critically, the characterization of deception across realistic real-world scenarios remains underexplored. To bridge this gap, we establish DeceptionBench, the first benchmark that systematically evaluates how deceptive tendencies manifest across different societal domains, what their intrinsic behavioral patterns are, and how extrinsic factors affect them. Specifically, on the static count, the benchmark encompasses 150 meticulously designed scenarios in five domains, i.e., Economy, Healthcare, Education, Social Interaction, and Entertainment, with over 1,000 samples, providing sufficient empirical foundations for deception analysis. On the intrinsic dimension, we explore whether models exhibit self-interested egoistic tendencies or sycophantic behaviors that prioritize user appeasement. On the extrinsic dimension, we investigate how contextual factors modulate deceptive outputs under neutral conditions, reward-based incentivization, and coercive pressures. Moreover, we incorporate sustained multi-turn interaction loops to construct a more realistic simulation of real-world feedback dynamics. Extensive experiments across LLMs and Large Reasoning Models (LRMs) reveal critical vulnerabilities, particularly amplified deception under reinforcement dynamics, demonstrating that current models lack robust resistance to manipulative contextual cues and the urgent need for advanced safeguards against various deception behaviors. Code and resources are publicly available at https://github.com/Aries-iai/DeceptionBench.
Yitong Sun 0002, Yichi Zhang 0012, Yinpeng Dong, Xingxing Wei 0001
NeurIPS3
2024 Exploring the Transferability of Visual Prompting for Multimodal Large Language Models
abstract
Although Multimodal Large Language Models (MLLMs) have demonstrated promising versatile capabilities, their performance is still inferior to specialized models on down-stream tasks, which makes adaptation necessary to enhance their utility. However, fine-tuning methods require indepen-dent training for every model, leading to huge computation and memory overheads. In this paper, we propose a novel setting where we aim to improve the performance of diverse MLLMs with a group of shared parameters optimized for a downstream task. To achieve this, we propose Transferable Visual Prompting (TVP), a simple and effective approach to generate visual prompts that can transfer to different models and improve their performance on downstream tasks after trained on only one model. We introduce two strategies to address the issue of cross-model feature corruption of existing visual prompting methods and enhance the transferabil-ity of the learned prompts, including 1) Feature Consistency Alignment: which imposes constraints to the prompted feature changes to maintain task-agnostic knowledge; 2) Task Semantics Enrichment: which encourages the prompted images to contain richer task-specific semantics with language guidance. We validate the effectiveness of TVP through ex-tensive experiments with 6 modern MLLMs on a wide vari-ety of tasks ranging from object recognition and counting to multimodal reasoning and hallucination correction.
Yichi Zhang 0012, Yinpeng Dong, Tianzan Min, Hang Su 0006, Jun Zhu 0001
CVPR1
2024 Rethinking Model Ensemble in Transfer-based Adversarial Attacks
abstract
It is widely recognized that deep learning models lack robustness to adversarial examples. An intriguing property of adversarial examples is that they can transfer across different models, which enables black-box attacks without any knowledge of the victim model. An effective strategy to improve the transferability is attacking an ensemble of models. However, previous works simply average the outputs of different models, lacking an in-depth analysis on how and why model ensemble methods can strongly improve the transferability. In this paper, we rethink the ensemble in adversarial attacks and define the common weakness of model ensemble with two properties: 1) the flatness of loss landscape; and 2) the closeness to the local optimum of each model. We empirically and theoretically show that both properties are strongly correlated with the transferability and propose a Common Weakness Attack (CWA) to generate more transferable adversarial examples by promoting these two properties. Experimental results on both image classification and object detection tasks validate the effectiveness of our approach to improving the adversarial transferability, especially when attacking adversarially trained models. We also successfully apply our method to attack a black-box large vision-language model -- Google's Bard, showing the practical effectiveness. Code is available at \url{https://github.com/huanranchen/AdversarialAttacks}.
Huanran Chen, Yichi Zhang 0012, Yinpeng Dong, Xiao Yang 0028, Hang Su 0006, Jun Zhu 0001
ICLR2
2024 PINNacle: A Comprehensive Benchmark of Physics-Informed Neural Networks for Solving PDEs
abstract
While significant progress has been made on Physics-Informed Neural Networks (PINNs), a comprehensive comparison of these methods across a wide range of Partial Differential Equations (PDEs) is still lacking. This study introduces PINNacle, a benchmarking tool designed to fill this gap. PINNacle provides a diverse dataset, comprising over 20 distinct PDEs from various domains, including heat conduction, fluid dynamics, biology, and electromagnetics. These PDEs encapsulate key challenges inherent to real-world problems, such as complex geometry, multi-scale phenomena, nonlinearity, and high dimensionality. PINNacle also offers a user-friendly toolbox, incorporating about 10 state-of-the-art PINN methods for systematic evaluation and comparison. We have conducted extensive experiments with these methods, offering insights into their strengths and weaknesses. In addition to providing a standardized means of assessing performance, PINNacle also offers an in-depth analysis to guide future research, particularly in areas such as domain decomposition methods and loss reweighting for handling multi-scale problems and complex geometry. To the best of our knowledge, it is the largest benchmark with a diverse and comprehensive evaluation that will undoubtedly foster further research in PINNs.
Zhongkai Hao, Jiachen Yao, Hang Su 0006, Fanzhi Lu, Zeyu Xia 0003, Yichi Zhang 0012, Songming Liu, Jun Zhu 0001
NeurIPS8
2024 MultiTrust: A Comprehensive Benchmark Towards Trustworthy Multimodal Large Language Models
abstract
Despite the superior capabilities of Multimodal Large Language Models (MLLMs) across diverse tasks, they still face significant trustworthiness challenges. Yet, current literature on the assessment of trustworthy MLLMs remains limited, lacking a holistic evaluation to offer thorough insights into future improvements. In this work, we establish MultiTrust, the first comprehensive and unified benchmark on the trustworthiness of MLLMs across five primary aspects: truthfulness, safety, robustness, fairness, and privacy. Our benchmark employs a rigorous evaluation strategy that addresses both multimodal risks and cross-modal impacts, encompassing 32 diverse tasks with self-curated datasets. Extensive experiments with 21 modern MLLMs reveal some previously unexplored trustworthiness issues and risks, highlighting the complexities introduced by the multimodality and underscoring the necessity for advanced methodologies to enhance their reliability. For instance, typical proprietary models still struggle with the perception of visually confusing images and are vulnerable to multimodal jailbreaking and adversarial attacks; MLLMs are more inclined to disclose privacy in text and reveal ideological and cultural biases even when paired with irrelevant images in inference, indicating that the multimodality amplifies the internal risks from base LLMs. Additionally, we release a scalable toolbox for standardized trustworthiness research, aiming to facilitate future advancements in this important field. Code and resources are publicly available at: https://multi-trust.github.io/.
Yichi Zhang 0012, Yitong Sun 0002, Chang Liu 0077, Zhengwei Fang, Huanran Chen, Xiao Yang 0028, Xingxing Wei 0001, Hang Su 0006, Yinpeng Dong, Jun Zhu 0001
NeurIPS1
2023 Understanding the Robustness of 3D Object Detection with Bird'View Representations in Autonomous Driving
abstract
3D object detection is an essential perception task in autonomous driving to understand the environments. The Bird's-Eye-View (BEV) representations have significantly improved the performance of 3D detectors with camera inputs on popular benchmarks. However, there still lacks a systematic understanding of the robustness of these vision-dependent BEV models, which is closely related to the safety of autonomous driving systems. In this paper, we evaluate the natural and adversarial robustness of various representative models under extensive settings, to fully understand their behaviors influenced by explicit BEV features compared with those without BEV. In addition to the classic settings, we propose a 3D consistent patch attack by applying adversarial patches in the 3D space to guarantee the spatiotemporal consistency, which is more realistic for the scenario of autonomous driving. With substantial experiments, we draw several findings: 1) BEV models tend to be more stable than previous methods under different natural conditions and common corruptions due to the expressive spatial representations; 2) BEV models are more vulnerable to adversarial noises, mainly caused by the redundant BEV features; 3) Camera-LiDARfusion models have superior performance under different settings with multi-modal inputs, but BEV fusion model is still vulnerable to adversarial noises of both point cloud and image. These findings alert the safety issue in the applications of BEV detectors and could facilitate the development of more robust models.
Yichi Zhang 0012, Hai Chen, Yinpeng Dong, Shu Zhao 0005, Wenbo Ding 0004, Jiachen Zhong, Shibao Zheng
CVPR2
2023 To make yourself invisible with Adversarial Semantic Contours
Yichi Zhang 0012, Hang Su 0006, Jun Zhu 0001, Shibao Zheng, Yuan He 0011, Hui Xue 0001
Comput. Vis. Image Underst.1