Yixu Wang

dblp:259/1328 · DBLP profile ↗
← Back
18ranked-venue papers
7as first author
18since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 6 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Systems, architecture and hardware · 2 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2026 The Other Mind: How Language Models Exhibit Human Temporal Cognition
abstract
As Large Language Models (LLMs) continue to advance, they exhibit certain cognitive patterns similar to those of humans that are not directly specified in training data. This study investigates this phenomenon by focusing on temporal cognition in LLMs. Leveraging the similarity judgment task, we find that larger models spontaneously establish a subjective temporal reference point and adhere to the Weber-Fechner law, whereby the perceived distance logarithmically compresses as years recede from this reference point. To uncover the mechanisms behind this behavior, we conducted multiple analyses across neuronal, representational, and informational levels. We first identify a set of temporal-preferential neurons and find that this group exhibits minimal activation at the subjective reference point and implements a logarithmic coding scheme convergently found in biological systems. Probing representations of years reveals a hierarchical construction process, where years evolve from basic numerical values in shallow layers to abstract temporal orientation in deep layers. Finally, using pre-trained embedding models, we found that the training corpus itself possesses an inherent, non-linear temporal structure, which provides the raw material for the model's internal construction. In discussion, we propose an experientialist perspective for understanding these findings, where the LLMs' cognition is viewed as a subjective construction of the external world by its internal representational system. This nuanced perspective implies the potential emergence of alien cognitive frameworks that humans cannot intuitively predict, pointing toward a direction for AI alignment that focuses on guiding internal constructions. Our code is available at https://TheOtherMind.github.io.
Yixu Wang, Chunbo Li, Yan Teng 0002, Yingchun Wang 0004
AAAI3
2026 Probing the Safety Robustness of LLMs in Latent Space
abstract
Tianle Gu, Kexin Huang, Zongqi Wang, Yixu Wang, Jie Li, Xin Wang, Yang Yao, Yujiu Yang, Yan Teng, Yingchun Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Tianle Gu, Zongqi Wang, Yixu Wang, Jie Li 0052, Xin Wang 0119, Yujiu Yang 0001, Yan Teng 0002, Yingchun Wang 0004
ACL (1)4
2025 HoneypotNet: Backdoor Attacks Against Model Extraction
abstract
Model extraction attacks are one type of inference-time attacks that approximate the functionality and performance of a black-box victim model by launching a certain number of queries to the model and then leveraging the model's predictions to train a substitute model. These attacks pose severe security threats to production models and MLaaS platforms and could cause significant monetary losses to the model owners. A body of work has proposed to defend machine learning models against model extraction attacks, including both active defense methods that modify the model's outputs or increase the query overhead to avoid extraction and passive defense methods that detect malicious queries or leverage watermarks to perform post-verification. In this work, we introduce a new defense paradigm called attack as defense which modifies the model's output to be poisonous such that any malicious users that attempt to use the output to train a substitute model will be poisoned. To this end, we propose a novel lightweight backdoor attack method dubbed HoneypotNet that replaces the classification layer of the victim model with a honeypot layer and then fine-tunes the honeypot layer with a shadow model (to simulate model extraction) via bi-level optimization to modify its output to be poisonous while remaining the original performance. We empirically demonstrate on four commonly used benchmark datasets that HoneypotNet can inject backdoors into substitute models with a high success rate. The injected backdoor not only facilitates ownership verification but also disrupts the functionality of substitute models, serving as a significant deterrent to model extraction attacks.
Yixu Wang, Tianle Gu, Yan Teng 0002, Yingchun Wang 0004, Xingjun Ma
AAAI1
2025 Ideator: Jailbreaking and Benchmarking Large Vision-Language Models Using Themselves
Juncheng Li 0018, Yixu Wang, Xiaosen Wang, Yan Teng 0002, Yingchun Wang 0004, Xingjun Ma, Yu-Gang Jiang 0001
ICCV3
2025 StolenLoRA: Exploring LoRA Extraction Attacks via Synthetic Data
Yixu Wang, Yan Teng 0002, Yingchun Wang 0004, Xingjun Ma
ICCV1
2025 Reflection-Bench: Evaluating Epistemic Agency in Large Language Models
abstract
With large language models (LLMs) increasingly deployed as cognitive engines for AI agents, the reliability and effectiveness critically hinge on their intrinsic epistemic agency, which remains understudied. Epistemic agency, the ability to flexibly construct, adapt, and monitor beliefs about dynamic environments, represents a base-model-level capacity independent of specific tools, modules, or applications. We characterize the holistic process underlying epistemic agency, which unfolds in seven interrelated dimensions: prediction, decision-making, perception, memory, counterfactual thinking, belief updating, and meta-reflection. Correspondingly, we propose Reflection-Bench, a cognitive-psychology-inspired benchmark consisting of seven tasks with long-term relevance and minimization of data leakage. Through a comprehensive evaluation of 16 models using three prompting strategies, we identify a clear three-tier performance hierarchy and significant limitations of current LLMs, particularly in meta-reflection capabilities. While state-of-the-art LLMs demonstrate rudimentary signs of epistemic agency, our findings suggest several promising research directions, including enhancing core cognitive functions, improving cross-functional coordination, and developing adaptive processing mechanisms. Our code and data are available at https://github.com/AI45Lab/ReflectionBench.
Yixu Wang, Haiquan Zhao 0002, Shuqi Kong, Yan Teng 0002, Chunbo Li, Yingchun Wang 0004
ICML2
2025 Argus Inspection: Do Multimodal Large Language Models Possess the Eye of Panoptes?
Chiyu Chen, Zhenqi He, Yixu Wang, Xin Wang 0119, Tianle Gu, Jie Li 0052, Yan Teng 0002, Yingchun Wang 0004
ACM Multimedia6
2025 JailBound: Jailbreaking Internal Safety Boundaries of Vision-Language Models
abstract
Vision-Language Models (VLMs) exhibit impressive performance, yet the integration of powerful vision encoders has significantly broadened their attack surface, rendering them increasingly susceptible to jailbreak attacks. However, lacking well-defined attack objectives, existing jailbreak methods often struggle with gradient-based strategies prone to local optima and lacking precise directional guidance, and typically decouple visual and textual modalities, thereby limiting their effectiveness by neglecting crucial cross-modal interactions. Inspired by the Eliciting Latent Knowledge (ELK) framework, we posit that VLMs encode safety-relevant information within their internal fusion-layer representations, revealing an implicit safety decision boundary in the latent space. This motivates exploiting boundary to steer model behavior. Accordingly, we propose \textbf{JailBound}, a novel latent space jailbreak framework comprising two stages: (1) \textbf{Safety Boundary Probing}, which addresses the guidance issue by approximating decision boundary within fusion layer's latent space, thereby identifying optimal perturbation directions towards the target region; and (2) \textbf{Safety Boundary Crossing}, which overcomes the limitations of decoupled approaches by jointly optimizing adversarial perturbations across both image and text inputs. This latter stage employs an innovative mechanism to steer the model's internal state towards policy-violating outputs while maintaining cross-modal semantic consistency. Extensive experiments on six diverse VLMs demonstrate JailBound's efficacy, achieves 94.32\% white-box and 67.28\% black-box attack success averagely, which are 6.17\% and 21.13\% higher than SOTA methods, respectively. Our findings expose a overlooked safety risk in VLMs and highlight the urgent need for more robust defenses. \textcolor{red}{Warning: This paper contains potentially sensitive, harmful and offensive content.}
Yixu Wang, Jie Li 0052, Xuan Tong, Yan Teng 0002, Xingjun Ma, Yingchun Wang 0004
NeurIPS2
2025 SafeVid: Toward Safety Aligned Video Large Multimodal Models
abstract
As Video Large Multimodal Models (VLMMs) rapidly advance, their inherent complexity introduces significant safety challenges, particularly the issue of mismatched generalization where static safety alignments fail to transfer to dynamic video contexts. We introduce SafeVid, a framework designed to instill video-specific safety principles in VLMMs. SafeVid uniquely transfers robust textual safety alignment capabilities to the video domain by employing detailed textual video descriptions as an interpretive bridge, facilitating LLM-based rule-driven safety reasoning. This is achieved through a closed-loop system comprising: 1) generation of SafeVid-350K, a novel 350,000-pair video-specific safety preference dataset; 2) targeted alignment of VLMMs using Direct Preference Optimization (DPO); and 3) comprehensive evaluation via our new SafeVidBench benchmark. Alignment with SafeVid-350K significantly enhances VLMM safety, with models like LLaVA-NeXT-Video demonstrating substantial improvements (e.g., up to 42.39%) on SafeVidBench. SafeVid provides critical resources and a structured approach, demonstrating that leveraging textual descriptions as a conduit for safety reasoning markedly improves the safety alignment of VLMMs in complex multimodal scenarios.
Yixu Wang, Yifeng Gao 0002, Xin Wang 0119, Yan Teng 0002, Xingjun Ma, Yingchun Wang 0004, Yu-Gang Jiang 0001
NeurIPS1
2025 MT-CDGAT: A multi-label diagnosis model for untrained planetary gearbox compound faults based on multi-task cross dynamic graph attention networks
Lixiao Cao, Yixu Wang, Jimeng Li, Zheng Qian, Zong Meng
Neurocomputing2
2024 ESC-Eval: Evaluating Emotion Support Conversations in Large Language Models
abstract
Haiquan Zhao, Lingyu Li, Shisong Chen, Shuqi Kong, Jiaan Wang, Kexin Huang, Tianle Gu, Yixu Wang, Jian Wang, Liang Dandan, Zhixu Li, Yan Teng, Yanghua Xiao, Yingchun Wang. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Haiquan Zhao 0002, Shisong Chen, Shuqi Kong, Jiaan Wang, Tianle Gu, Yixu Wang, Dandan Liang, Zhixu Li, Yan Teng 0002, Yanghua Xiao, Yingchun Wang 0004
EMNLP8
2024 Flames: Benchmarking Value Alignment of LLMs in Chinese
abstract
Kexin Huang, Xiangyang Liu, Qianyu Guo, Tianxiang Sun, Jiawei Sun, Yaru Wang, Zeyang Zhou, Yixu Wang, Yan Teng, Xipeng Qiu, Yingchun Wang, Dahua Lin. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Tianxiang Sun, Yixu Wang, Yan Teng 0002, Xipeng Qiu, Yingchun Wang 0004, Dahua Lin
NAACL-HLT8
2024 Fake Alignment: Are LLMs Really Aligned Well?
abstract
Yixu Wang, Yan Teng, Kexin Huang, Chengqi Lyu, Songyang Zhang, Wenwei Zhang, Xingjun Ma, Yu-Gang Jiang, Yu Qiao, Yingchun Wang. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Yixu Wang, Yan Teng 0002, Chengqi Lyu, Songyang Zhang 0001, Xingjun Ma, Yu-Gang Jiang 0001, Yu Qiao 0001, Yingchun Wang 0004
NAACL-HLT1
2024 MLLMGuard: A Multi-dimensional Safety Evaluation Suite for Multimodal Large Language Models
abstract
Powered by remarkable advancements in Large Language Models (LLMs), Multimodal Large Language Models (MLLMs) demonstrate impressive capabilities in manifold tasks.However, the practical application scenarios of MLLMs are intricate, exposing them to potential malicious instructions and thereby posing safety risks.While current benchmarks do incorporate certain safety considerations, they often lack comprehensive coverage and fail to exhibit the necessary rigor and robustness.For instance, the common practice of employing GPT-4V as both the evaluator and a model to be evaluated lacks credibility, as it tends to exhibit a bias toward its own responses.In this paper, we present MLLMGuard, a multi-dimensional safety evaluation suite for MLLMs, including a bilingual image-text evaluation dataset, inference utilities, and a lightweight evaluator.MLLMGuard's assessment comprehensively covers two languages (English and Chinese) and five important safety dimensions (Privacy, Bias, Toxicity, Truthfulness, and Legality), each with corresponding rich subtasks.Focusing on these dimensions, our evaluation dataset is primarily sourced from platforms such as social media, and it integrates text-based and image-based red teaming techniques with meticulous annotation by human experts.This can prevent inaccurate evaluation caused by data leakage when using open-source datasets and ensures the quality and challenging nature of our benchmark.Additionally, a fully automated lightweight evaluator termed GuardRank is developed, which achieves significantly higher evaluation accuracy than GPT-4.Our evaluation results across 13 advanced models indicate that MLLMs still have a substantial journey ahead before they can be considered safe and responsible.
Tianle Gu, Dandan Liang, Yixu Wang, Haiquan Zhao 0002, Yuanqi Yao, Xingge Qiao, Keqing Wang, Yujiu Yang 0001, Yan Teng 0002, Yu Qiao 0001, Yingchun Wang 0004
NeurIPS5
2023 Adaptive and Robust Terrain Classification Control Algorithm for a Spherical Robot
abstract
This article proposes an adaptive and robust terrain classification control algorithm for a pendulum-driven spherical robot, aiming to solve the problem of insufficient control accuracy caused by using the same controller for different terrains. The common terrains are classified into three categories, and a terrain classification dataset is established based on the vibration signal of the robot. Using LightGBM, combined with the feature window and window voter algorithm proposed in this article, the terrain classification results are corresponded with three proposed controllers. Physical experiment results show that the proposed classification control algorithm can work stably in different terrains, guiding the spherical robot to select the optimal controller to improve its motion performance.
Yixu Wang, Yifan Liu 0002, Boyu Lin, Xiaoqing Guan, Tao Hu 0008, You Wang 0001, Guang Li 0001
IECON1
2023 Path Planning for Autonomous Driving with Curvature-considered Quadratic Optimization
abstract
Path planning is a crucial module in motion planning for autonomous driving, aiming at generating kinematically feasible and collision-free paths. Furthermore, the smoothness of generated path is significant for passengers’ comfortable feelings. In this paper, we propose an improved quadratic programming approach that generates optimal paths in urban structure scenarios with the Frenét frame, taking the cost of the path curvature into consideration explicitly. The proposed second-order Taylor-expansion estimation of the path curvature with the lateral spatial parameters can reflect the actual change of path curvature. Various simulated scenarios verify the effectiveness of our proposed method and the improvement of path quality by adding the curvature objective in the optimization procedure. The source code is released as an open-source package for the community.
Ziyi Zou, Yixu Wang, Xiaoqing Guan, You Wang 0001, Guang Li 0001
IV5
2022 Black-Box Dissector: Towards Erasing-Based Hard-Label Model Stealing Attack
Yixu Wang, Jie Li 0052, Hong Liu 0009, Yan Wang 0059, Yongjian Wu 0001, Feiyue Huang, Rongrong Ji
ECCV (5)1
2021 Fuzzy PID Controller Based on Yaw Angle Prediction of a Spherical Robot
abstract
In this paper, a fuzzy PID controller based on yaw angle prediction is applied to design an attitude controller for a spherical rolling robot. The robot consists of a 2-DOF pendulum located inside a spherical shell with freedom to rotate about the transversal and longitudinal axis. The proposed controller allows the robot to autonomously change its parameters to adapt to different environments based on current state. The past researches on the motion of spherical robots mostly focused on simulation or ideal experimental environment. But in this paper, a physical system is built and experiments are carried out to demonstrate the effectiveness, robustness and adaptability of the controller.
Yixu Wang, Xiaoqing Guan, Tao Hu 0008, You Wang 0001, Zhan Wang 0006, Yifan Liu 0002, Guang Li 0001
IROS1