Biao Yi

dblp:290/6528 · DBLP profile ↗
← Back
13ranked-venue papers
5as first author
13since 2021 · last 2026
0000-0002-8347-1953ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author · 4 since 2021Security and privacy · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 EcoAgent: An Efficient Device-Cloud Collaborative Multi-Agent Framework for Mobile Automation
abstract
To tackle increasingly complex tasks, recent research on mobile agents has shifted towards multi-agent collaboration. Current mobile multi-agent systems are primarily deployed in the cloud, leading to high latency and operational costs. A straightforward idea is to deploy a device–cloud collaborative multi-agent system, which is nontrivial, as directly extending existing systems introduces new challenges: (1) reliance on cloud-side verification requires uploading mobile screenshots, compromising user privacy; and (2) open-loop cooperation lacking device-to-cloud feedback, underutilizing device resources and increasing latency. To overcome these limitations, we propose EcoAgent, a closed-loop device-cloud collaborative multi-agent framework designed for privacy-aware, efficient, and responsive mobile automation. EcoAgent integrates a novel reasoning approach, Dual-ReACT, into the cloud-based Planning Agent, fully exploiting cloud reasoning to compensate for limited on-device capacity, thereby enabling device-side verification and lightweight feedback. Furthermore, the device-based Observation Agent leverages a Pre-understanding Module to summarize screen content into concise textual descriptions, significantly reducing token usage and device-cloud communication overhead while preserving privacy. Experiments on AndroidWorld demonstrate that EcoAgent matches the task success rates of fully cloud-based agents, while reducing resource consumption and response latency.
Biao Yi, Xueyu Hu, Yurun Chen 0004, Shengyu Zhang 0001, Hongxia Yang, Fan Wu 0006
AAAI1
2026 CTRAP: Embedding Collapse Trap to Safeguard Large Language Models from Harmful Fine-Tuning
abstract
Fine-tuning-as-a-service, while commercially successful for Large Language Model (LLM) providers, exposes models to harmful finetuning attacks.As a widely explored defense paradigm against such attacks, unlearning attempts to remove malicious knowledge from LLMs, thereby essentially preventing them from being used to perform malicious tasks.However, we highlight a critical flaw: the inherent general adaptability of LLMs allows them to easily bypass selective unlearning by rapidly relearning or repurposing their general capabilities for harmful tasks.To address this fundamental limitation, we propose a paradigm shift: instead of selective removal, we advocate for inducing model collapse, effectively forcing the model to "unlearn everything", specifically in response to updates characteristic of malicious adaptation.This collapse directly neutralizes the very general capabilities that attackers exploit, tackling the core issue unaddressed by selective unlearning.We introduce the Collapse Trap (CTRAP) as a practical mechanism to implement this concept conditionally.Embedded during alignment, CTRAP pre-configures the model's reaction to subsequent fine-tuning dynamics.If updates during fine-tuning constitute a persistent attempt to reverse safety alignment, the pre-configured trap triggers a progressive degradation of the model's core language modeling abilities, ultimately rendering it inert and useless for the attacker.Crucially, this collapse mechanism remains dormant during benign fine-tuning, ensuring the model's utility and general capabilities are preserved.1
Biao Yi, Tiansheng Huang, Baolei Zhang, Tong Li 0011, Lihai Nie, Zheli Liu, Li Shen 0008
ACL (1)1
2026 Who Taught the Lie? Responsibility Attribution for Poisoned Knowledge in Retrieval-Augmented Generation
abstract
Retrieval-Augmented Generation (RAG) integrates external knowledge into large language models to improve response quality. However, recent work has shown that RAG systems are highly vulnerable to poisoning attacks, where malicious texts are inserted into the knowledge database to influence model outputs. While several defenses have been proposed, they are often circumvented by more adaptive or sophisticated attacks. This paper presents RAGOrigin, a black-box responsibility attribution framework designed to identify which texts in the knowledge database are responsible for misleading or incorrect generations. Our method constructs a focused attribution scope tailored to each misgeneration event and assigns a responsibility score to each candidate text by evaluating its retrieval ranking, semantic relevance, and influence on the generated response. The system then isolates poisoned texts using an unsupervised clustering method. We evaluate RAGOrigin across seven datasets and fifteen poisoning attacks, including newly developed adaptive poisoning strategies and multi-attacker scenarios. Our approach outperforms existing baselines in identifying poisoned content and remains robust under dynamic and noisy conditions. These results suggest that RAGOrigin provides a practical and effective solution for tracing the origins of corrupted knowledge in RAG systems. Our code is available at: https://github.com/zhangbl6618/RAG-Responsibility-Attribution
Baolei Zhang, Haoran Xin 0002, Zhuqing Liu, Biao Yi, Tong Li 0011, Lihai Nie, Zheli Liu, Minghong Fang
SP5
2026 Practical Framework for Privacy-Preserving and Byzantine-Robust Federated Learning
abstract
Federated Learning (FL) allows multiple clients to collaboratively train a model without sharing their private data. However, FL is vulnerable toByzantine attacks, where adversaries manipulate client models to compromise the federated model, andprivacy inference attacks, where adversaries exploit client models to infer private data. Existing defenses against both backdoor and privacy inference attacks introduce significant computational and communication overhead, creating a gap between theory and practice. To address this, we propose ABBR, a practical framework for Byzantine-robust and privacy-preserving FL. We are the first to utilize dimensionality reduction to speed up the private computation of complex filtering rules in privacy-preserving FL. Additionally, we analyze the accuracy loss of vector-wise filtering in low-dimensional space and introduce an adaptive tuning strategy to minimize the impact of malicious models that bypass filtering on the global model. We implement ABBR with state-of-the-art Byzantine-robust aggregation rules and evaluate it on public datasets, showing that it runs significantly faster, has minimal communication overhead, and maintains nearly the same Byzantine-resilience as the baselines.
Baolei Zhang, Minghong Fang, Zhuqing Liu, Biao Yi, Peizhao Zhou, Tong Li 0011, Zheli Liu
IEEE Trans. Inf. Forensics Secur.4
2025 OS Agents: A Survey on MLLM-based Agents for Computer, Phone and Browser Use
abstract
Xueyu Hu, Tao Xiong, Biao Yi, Zishu Wei, Ruixuan Xiao, Yurun Chen, Jiasheng Ye, Meiling Tao, Xiangxin Zhou, Ziyu Zhao, Yuhuai Li, Shengze Xu, Shenzhi Wang, Xinchen Xu, Shuofei Qiao, Zhaokai Wang, Kun Kuang, Tieyong Zeng, Liang Wang, Jiwei Li, Yuchen Eleanor Jiang, Wangchunshu Zhou, Guoyin Wang, Keting Yin, Zhou Zhao, Hongxia Yang, Fan Wu, Shengyu Zhang, Fei Wu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Xueyu Hu, Biao Yi, Zishu Wei, Ruixuan Xiao, Yurun Chen 0004, Jiasheng Ye, Meiling Tao, Xiangxin Zhou, Ziyu Zhao 0001, Yuhuai Li, Shengze Xu, Shenzhi Wang, Shuofei Qiao, Zhaokai Wang, Kun Kuang 0001, Tieyong Zeng, Liang Wang 0001, Jiwei Li 0001, Yuchen Eleanor Jiang, Wangchunshu Zhou, Guoyin Wang 0002, Keting Yin, Zhou Zhao 0001, Hongxia Yang, Fan Wu 0006, Shengyu Zhang 0001, Fei Wu 0001
ACL (1)3
2025 Prompt-Guided Internal States for Hallucination Detection of Large Language Models
abstract
Large Language Models (LLMs) have demonstrated remarkable capabilities across a variety of tasks in different domains. However, they sometimes generate responses that are logically coherent but factually incorrect or misleading, which is known as LLM hallucinations. Data-driven supervised methods train hallucination detectors by leveraging the internal states of LLMs, but detectors trained on specific domains often struggle to generalize well to other domains. In this paper, we aim to enhance the cross-domain performance of supervised detectors with only in-domain data. We propose a novel framework, prompt-guided internal states for hallucination detection of LLMs, namely PRISM. By utilizing appropriate prompts to guide changes to the structure related to text truthfulness in LLMs’ internal states, we make this structure more salient and consistent across texts from different domains. We integrated our framework with existing hallucination detection methods and conducted experiments on datasets from different domains. The experimental results indicate that our framework significantly enhances the cross-domain generalization of existing hallucination detection methods.
Fujie Zhang, Peiqi Yu, Biao Yi, Baolei Zhang, Tong Li 0011, Zheli Liu
ACL (1)3
2025 Probe before You Talk: Towards Black-box Defense against Backdoor Unalignment for Large Language Models
abstract
Backdoor unalignment attacks against Large Language Models (LLMs) enable the stealthy compromise of safety alignment using a hidden trigger while evading normal safety auditing. These attacks pose significant threats to the applications of LLMs in the real-world Large Language Model as a Service (LLMaaS) setting, where the deployed model is a fully black-box system that can only interact through text. Furthermore, the sample-dependent nature of the attack target exacerbates the threat. Instead of outputting a fixed label, the backdoored LLM follows the semantics of any malicious command with the hidden trigger, significantly expanding the target space. In this paper, we introduce BEAT, a black-box defense that detects triggered samples during inference to deactivate the backdoor. It is motivated by an intriguing observation (dubbed the **probe concatenate effect**), where concatenated triggered samples significantly reduce the refusal rate of the backdoored LLM towards a malicious probe, while non-triggered samples have little effect. Specifically, BEAT identifies whether an input is triggered by measuring the degree of distortion in the output distribution of the probe before and after concatenation with the input. Our method addresses the challenges of sample-dependent targets from an opposite perspective. It captures the impact of the trigger on the refusal signal (which is sample-independent) instead of sample-specific successful attack behaviors. It overcomes black-box access limitations by using multiple sampling to approximate the output distribution. Extensive experiments are conducted on various backdoor attacks and LLMs (including the closed-source GPT-3.5-turbo), verifying the effectiveness and efficiency of our defense. Besides, we also preliminarily verify that BEAT can effectively defend against popular jailbreak attacks, as they can be regarded as "natural backdoors". Our source code is available at https://github.com/clearloveclearlove/BEAT.
Biao Yi, Tiansheng Huang, Sishuo Chen, Tong Li 0011, Zheli Liu, Zhixuan Chu, Yiming Li 0004
ICLR1
2025 Traceback of Poisoning Attacks to Retrieval-Augmented Generation
abstract
Large language models (LLMs) integrated with retrieval-augmented generation (RAG) systems improve accuracy by leveraging external knowledge sources. However, recent research has revealed RAG's susceptibility to poisoning attacks, where the attacker injects poisoned texts into the knowledge database, leading to attacker-desired responses. Existing defenses, which predominantly focus on inference-time mitigation, have proven insufficient against sophisticated attacks. In this paper, we introduce RAGForensics, the first traceback system for RAG, designed to identify poisoned texts within the knowledge database that are responsible for the attacks. RAGForensics operates iteratively, first retrieving a subset of texts from the database and then utilizing a specially crafted prompt to guide an LLM in detecting potential poisoning texts. Empirical evaluations across multiple datasets demonstrate the effectiveness of RAGForensics against state-of-the-art poisoning attacks. This work pioneers the traceback of poisoned texts in RAG systems, providing a practical and promising defense mechanism to enhance their security.
Baolei Zhang, Haoran Xin 0002, Minghong Fang, Zhuqing Liu, Biao Yi, Tong Li 0011, Zheli Liu
WWW5
2024 Semantic-Preserving Linguistic Steganography by Pivot Translation and Semantic-Aware Bins Coding
abstract
Linguistic steganography (LS) aims to embed secret information into a highly encoded text for covert communication. It can be roughly divided to two main categories, i.e., modification based LS (MLS) and generation based LS (GLS). MLS embeds secret data by slightly modifying a given text without impairing the meaning of the text, whereas GLS uses a well trained language model to directly generate a text carrying secret data. A common disadvantage for MLS methods is that the embedding payload is very small, whose return is well preserving the semantic quality of the text. In contrast, GLS enables the data hider to embed a large payload, which has to pay the high price of uncontrollable semantics. In this article, we propose a novel LS method to modify a given text by pivoting it between two different languages and embed secret data using a semantic-aware information encoding strategy. Our purpose is to alter the expression of the given text, enabling a large payload to be embedded while keeping the semantic information unchanged. Experiments have shown that the proposed work not only achieves a large embedding payload, but also shows superior performance in maintaining the semantic consistency and resisting linguistic steganalysis.
Hanzhou Wu, Biao Yi, Guorui Feng, Xinpeng Zhang 0001
IEEE Trans. Dependable Secur. Comput.3
2022 Exploiting Language Model For Efficient Linguistic Steganalysis
abstract
Recent advances in linguistic steganalysis have successively applied CNN, RNN, GNN and other efficient deep models for detecting secret information in generative texts. These methods tend to seek stronger feature extractors to achieve higher steganalysis effects. However, we have found through experiments that there actually exists significant difference between automatically generated stego texts and carrier texts in terms of the conditional probability distribution of individual words. Such kind of difference can be naturally captured by the language model used for generating stego texts. Through further experiments, we conclude that this ability can be transplanted to a text classifier by pre-training and fine-tuning to improve the detection performance. Motivated by this insight, we propose two methods for efficient linguistic steganalysis. One is to pre-train a language model based on RNN, and the other is to pre-train a sequence autoencoder. The results indicate that the two methods have different degrees of performance gain compared to the randomly initialized RNN, and the convergence speed is significantly accelerated. Moreover, our methods achieved the best performance compared to related works, while providing a solution for real-world scenario where there are more cover texts than stego texts.
Biao Yi, Hanzhou Wu, Guorui Feng, Xinpeng Zhang 0001
ICASSP1
2022 Link Prediction via Fused Attribute Features Activation with Graph Convolutional Network
Yayao Zuo, Biao Yi, Minghao Zhan
PRICAI (2)3
2022 ALiSa: Acrostic Linguistic Steganography Based on BERT and Gibbs Sampling
abstract
In this letter, we propose a novel linguistic steganographic method that directly conceals a token-level secret message in a seemingly-natural steganographic text generated by the off-the-shelf BERT model equipped with Gibbs sampling. Compared with all modification based linguistic steganographic methods, the proposed method does not modify a given cover text. Instead, the proposed method utilizes the secret message to directly generate the steganographic text. Compared with mainstream generation based linguistic steganographic methods, the proposed method enables the receiver to collect the tokens of the specific positions to directly constitute the secret message, without a complex decoding process and much side information shared between the sender and the receiver. Experimental results show that the proposed method can generate fluent, highly readable steganographic texts, while enjoying pretty good anti-steganalysis ability. This work has great application potential in real-time covert communication.
Biao Yi, Hanzhou Wu, Guorui Feng, Xinpeng Zhang 0001
IEEE Signal Process. Lett.1
2021 Linguistic Steganalysis With Graph Neural Networks
abstract
Recent linguistic steganalysis methods model texts as sequences and use deep learning models to extract discriminative features for detecting the presence of secret information in texts. However, natural language has a complex syntactic structure and sequences have limited representation ability for text modeling. Moreover, previous methods tend to extract features from local continuous word sequences, which cannot effectively model global characteristics. In this paper, we present a linguistic steganalysis method with graph neural network. In the proposed method, texts are translated as directed graphs with the associated information, where nodes denote words and edges show associations between the words. By training a graph convolutional network for feature extraction, each node of a graph can collect contextual information to update self-expression, accordingly effectively solving the problem of poor representation of polysemous words. Meanwhile, we adopt a globally-shared matrix to record correlation strengths between words so that each text can effectively utilize the global information to obtain the better self-representation. Experimental results have shown that the proposed work achieves the state-of-the-art performance comparing with the previous works.
Hanzhou Wu, Biao Yi, Feng Ding 0007, Guorui Feng, Xinpeng Zhang 0001
IEEE Signal Process. Lett.2