Yuanpu Cao

dblp:243/0230 · DBLP profile ↗
← Back
13ranked-venue papers
3as first author
11since 2021 · last 2026
0009-0004-1993-912XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 3 first-author · 9 since 2021Computer networks · 2 · 1 since 2021Security and privacy · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
10 papers
Language models and text generation · 59% Efficient and distributed learning · 20% Generative modeling · 8%
Network and information security
5 papers
Security and privacy of machine learning · 100%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Medical and health informatics · 100%
Computer networks
1 paper
Network measurement and analytics · 100%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Performance modeling and evaluation · 77% Cloud and datacenter computing · 23%

Topics — the 26 heaviest of 29, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
hallucination mitigation
1.122025
TruthFlow: Truthful LLM Generation via Representation Flow Correction · ICML 2025
Personalized Steering of Large Language Models: Versatile Steering Vectors Through Bi-directional Preference Optimization · NeurIPS 2024
Natural language and speech › Language models and text generation
knowledge editing
1.012026
Can Factual Opinions Be Edited (Manipulated) in Large Language Models? · ACL (1) 2026
Natural language and speech › Language models and text generation
LLM agents
1.012026
ICDAGENT: Empowering Agentic Large Language Models for Explainable Medical Coding · ACL (1) 2026
Medical and health informatics › clinical informatics
clinical coding
1.012026
ICDAGENT: Empowering Agentic Large Language Models for Explainable Medical Coding · ACL (1) 2026
Machine learning › Generative modeling
diffusion model
0.912025
AdvI2I: Adversarial Image Attack on Image-to-Image Diffusion Models · ICML 2025
Security and privacy of machine learning
adversarial attack
0.912025
Phi: Preference Hijacking in Multi-modal Large Language Models at Inference Time · EMNLP 2025
Security and privacy of machine learning › adversarial attack
adversarial attacks on generative models
0.912025
AdvI2I: Adversarial Image Attack on Image-to-Image Diffusion Models · ICML 2025
Security and privacy of machine learning
large language model security
0.912025
You Can't Steal Nothing: Mitigating Prompt Leakages in LLMs via System Vectors · CCS 2025
Security and privacy of machine learning › large language model security
multimodal large language model security
0.912025
Phi: Preference Hijacking in Multi-modal Large Language Models at Inference Time · EMNLP 2025
Natural language and speech › Language models and text generation
alignment
0.812024
Personalized Steering of Large Language Models: Versatile Steering Vectors Through Bi-directional Preference Optimization · NeurIPS 2024
Machine learning › Efficient and distributed learning › federated learning
asynchronous federated learning
0.812024
Tackling the Data Heterogeneity in Asynchronous Federated Learning with Cached Update Calibration · ICLR 2024
Machine learning › Optimization for machine learning
convergence analysis
0.812024
Tackling the Data Heterogeneity in Asynchronous Federated Learning with Cached Update Calibration · ICLR 2024
Machine learning › Efficient and distributed learning › federated learning
data heterogeneity
0.812024
Tackling the Data Heterogeneity in Asynchronous Federated Learning with Cached Update Calibration · ICLR 2024
Machine learning › Efficient and distributed learning
federated learning
0.812024
Tackling the Data Heterogeneity in Asynchronous Federated Learning with Cached Update Calibration · ICLR 2024
Natural language and speech › Language models and text generation
preference optimization
0.812024
Personalized Steering of Large Language Models: Versatile Steering Vectors Through Bi-directional Preference Optimization · NeurIPS 2024
Natural language and speech › Language models and text generation
steering vectors
0.812024
Personalized Steering of Large Language Models: Versatile Steering Vectors Through Bi-directional Preference Optimization · NeurIPS 2024
Security and privacy of machine learning › large language model safety
jailbreak defense
0.812024
Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM · ACL (1) 2024
Security and privacy of machine learning
large language model alignment
0.812024
Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM · ACL (1) 2024
Network measurement and analytics
anomaly detection
0.512021
CTF: Anomaly Detection in High-Dimensional Time Series with Coarse-to-Fine Model Transfer · INFOCOM 2021
Network measurement and analytics › anomaly detection
time series anomaly detection
0.512021
CTF: Anomaly Detection in High-Dimensional Time Series with Coarse-to-Fine Model Transfer · INFOCOM 2021
Performance modeling and evaluation › performance monitoring
monitoring data analysis
0.512021
CTF: Anomaly Detection in High-Dimensional Time Series with Coarse-to-Fine Model Transfer · INFOCOM 2021
Knowledge, reasoning and agents › Multi-agent systems
imperfect information games
0.412020
RLCard: A Platform for Reinforcement Learning in Card Games · IJCAI 2020
Natural language and speech › Language models and text generation
instruction following
0.312025
You Can't Steal Nothing: Mitigating Prompt Leakages in LLMs via System Vectors · CCS 2025
Computer vision › Vision and language › vision-language model
multimodal large language model
0.312025
Phi: Preference Hijacking in Multi-modal Large Language Models at Inference Time · EMNLP 2025
Natural language and speech › Language models and text generation
large language model safety
0.212024
Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM · ACL (1) 2024
Cloud and datacenter computing › datacenter operations
datacenter monitoring
0.112021
CTF: Anomaly Detection in High-Dimensional Time Series with Coarse-to-Fine Model Transfer · INFOCOM 2021

Methods — techniques the papers use, named apart from their topics

self-generated evidence alignment · 2.0reinforcement learning · 2.0large language model · 2.0chain-of-thought reasoning · 2.0benchmark evaluation · 2.0universal perturbation · 1.7preference hijacking · 1.7adversarial perturbation · 1.7model transfer · 1.0fine-tuning · 1.0clustering · 1.0system prompt encoding · 0.9representation correction · 0.9internal representation encoding · 0.9generator training · 0.9flow matching · 0.9adversarial optimization · 0.9adversarial prompt analysis · 0.8
YearPublicationVenuePosition
2026 Can Factual Opinions Be Edited (Manipulated) in Large Language Models?
abstract
Large Language Models (LLMs) are increasingly integrated into various domains, making knowledge editing techniques crucial yet potentially hazardous.Current editing methods primarily target atomic facts, overlooking the significant risks associated with manipulating "factual opinions", e.g., documented stances of public figures on societal issues.Such manipulation could reshape public images, influence elections, and alter societal views.To systematically assess this threat, we introduce the Factual Opinion Editing with Evidence (FOE) benchmark, which encompasses 261 public figures, 19 issue categories, and 2,178 complete opinion records.Our evaluations demonstrate that current editing techniques struggle significantly with factual opinions, often achieving only superficial changes while failing to preserve consistency between the edited opinion and the supporting evidence generated by the model.To address this limitation, we further propose a simple yet effective Self-Generated Evidence-Aligned method that achieves opinion-evidence alignment without relying on explicit instructions.Together, our benchmark and method provide a foundation for understanding the emerging security implications of factual opinion editing in LLMs.
Yuanpu Cao, Ziyi Yin 0003, Fenglong Ma
ACL (1)1
2026 ICDAGENT: Empowering Agentic Large Language Models for Explainable Medical Coding
abstract
The explainable medical coding task aims to automatically assign International Classification of Diseases (ICD) codes to clinical notes while providing explicit justifications for each assignment.Recent approaches employ large language models (LLMs) to generate such explanations.However, their performance remains limited due to a lack of understanding of the clinical meanings of ICD codes.Additionally, the vast ICD code space further complicates the task of accurate prediction.To address these challenges, we propose the ICDAGENT framework, which consists of two collaborative LLM agents: a coding agent and a critical agent.The coding agent extracts ICD codes and generates preliminary rationales, while the critical agent performs fine-grained chain-of-thought reasoning to verify and refine them.Furthermore, the critical agent is trained with a rationaleaware reward, combined with reinforcement learning, enabling it to distinguish between correct and incorrect reasoning and ensure explanation accuracy.Experiments across multiple ICD coding standards and datasets demonstrate that ICDAGENT achieves effective ICD coding with accurate and trustworthy explanations.1
Ziyi Yin 0003, Yuanpu Cao, Ting Wang 0006, Fenglong Ma
ACL (1)2
2025 You Can't Steal Nothing: Mitigating Prompt Leakages in LLMs via System Vectors
abstract
Large language models (LLMs) have been widely adopted across various applications, leveraging customized system prompts for diverse tasks. Facing potential system prompt leakage risks, model developers have implemented strategies to prevent leakage, primarily by disabling LLMs from repeating their context when encountering known attack patterns. However, it remains vulnerable to new and unforeseen prompt-leaking techniques. In this paper, we first introduce a simple yet effective prompt leaking attack to reveal such risks. Our attack is capable of extracting system prompts from various LLM-based application, even from SOTA LLM models such as GPT-4o or Claude 3.5 Sonnet. Our findings further inspire us to search for a fundamental solution to the problems by having no system prompt in the context. To this end, we propose SysVec, a novel method that encodes system prompts as internal representation vectors rather than raw text. By doing so, SysVec minimizes the risk of unauthorized disclosure while preserving the LLM's core language capabilities. Remarkably, this approach not only enhances security but also improves the model's general instruction-following abilities. Experimental results demonstrate that SysVec effectively mitigates prompt leakage attacks, preserves the LLM's functional integrity, and helps alleviate the forgetting issue in long-context scenarios.
Bochuan Cao, Changjiang Li, Yuanpu Cao, Yameng Ge, Ting Wang 0006
CCS3
2025 Phi: Preference Hijacking in Multi-modal Large Language Models at Inference Time
abstract
Recently, Multi-modal Large Language Models (MLLMs) have gained significant attention across various domains.However, their widespread adoption has also raised serious safety concerns.In this paper, we uncover a new safety risk of MLLMs: the output preference of MLLMs can be arbitrarily manipulated by carefully optimized images.Such attacks often generate contextually relevant yet biased responses that are neither overtly harmful nor unethical, making them difficult to detect.Specifically, we introduce a novel method, Preference Hijacking (Phi), for manipulating the MLLM response preferences using a preference hijacked image.Our method works at inference time and requires no model modifications.Additionally, we introduce a universal hijacking perturbation -a transferable component that can be embedded into different images to hijack MLLM responses toward any attacker-specified preferences.Experimental results across various tasks demonstrate the effectiveness of our approach.The code for Phi is accessible at https://github.com/Yifan-Lan/Phi.
Yifan Lan, Yuanpu Cao, Lu Lin 0001
EMNLP2
2025 TruthFlow: Truthful LLM Generation via Representation Flow Correction
abstract
Large language models (LLMs) are known to struggle with consistently generating truthful responses. While various representation intervention techniques have been proposed, these methods typically apply a universal representation correction vector to all input queries, limiting their effectiveness against diverse queries in practice. In this study, we introduce TruthFlow, a novel method that leverages the Flow Matching technique for query-specific truthful representation correction. Specifically, TruthFlow first uses a flow model to learn query-specific correction vectors that transition representations from hallucinated to truthful states. Then, during inference, the trained flow model generates these correction vectors to enhance the truthfulness of LLM outputs. Experimental results demonstrate that TruthFlow significantly improves performance on open-ended generation tasks across various advanced LLMs evaluated on TruthfulQA. Moreover, the trained TruthFlow model exhibits strong transferability, performing effectively on other unseen hallucination benchmarks.
Bochuan Cao, Yuanpu Cao
ICML3
2025 AdvI2I: Adversarial Image Attack on Image-to-Image Diffusion Models
abstract
Recent advances in diffusion models have significantly enhanced the quality of image synthesis, yet they have also introduced serious safety concerns, particularly the generation of Not Safe for Work (NSFW) content. Previous research has demonstrated that adversarial prompts can be used to generate NSFW content. However, such adversarial text prompts are often easily detectable by text-based filters, limiting their efficacy. In this paper, we expose a previously overlooked vulnerability: adversarial image attacks targeting Image-to-Image (I2I) diffusion models. We propose AdvI2I, a novel framework that manipulates input images to induce diffusion models to generate NSFW content. By optimizing a generator to craft adversarial images, AdvI2I circumvents existing defense mechanisms, such as Safe Latent Diffusion (SLD), without altering the text prompts. Furthermore, we introduce AdvI2I-Adaptive, an enhanced version that adapts to potential countermeasures and minimizes the resemblance between adversarial images and NSFW concept embeddings, making the attack more resilient against defenses. Through extensive experiments, we demonstrate that both AdvI2I and AdvI2I-Adaptive can effectively bypass current safeguards, highlighting the urgent need for stronger security measures to address the misuse of I2I diffusion models.
Yaopei Zeng, Yuanpu Cao, Bochuan Cao, Yurui Chang, Lu Lin 0001
ICML2
2024 Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM
abstract
Recently, Large Language Models (LLMs) have made significant advancements and are now widely used across various domains.Unfortunately, there has been a rising concern that LLMs can be misused to generate harmful or malicious content.Though a line of research has focused on aligning LLMs with human values and preventing them from producing inappropriate content, such alignments are usually vulnerable and can be bypassed by alignmentbreaking attacks via adversarially optimized or handcrafted jailbreaking prompts.In this work, we introduce a Robustly Aligned LLM (RA-LLM) to defend against potential alignmentbreaking attacks.RA-LLM can be directly constructed upon an existing aligned LLM with a robust alignment checking function, without requiring any expensive retraining or fine-tuning process of the original LLM.Furthermore, we also provide a theoretical analysis for RA-LLM to verify its effectiveness in defending against alignment-breaking attacks.Through real-world experiments on open-source large language models, we demonstrate that RA-LLM can successfully defend against both state-of-the-art adversarial prompts and popular handcrafted jailbreaking prompts by reducing their attack success rates from nearly 100% to around 10% or less.WARNING: This paper contains unsafe model responses.Reader discretion is advised.
Bochuan Cao, Yuanpu Cao, Lu Lin 0001
ACL (1)2
2024 Tackling the Data Heterogeneity in Asynchronous Federated Learning with Cached Update Calibration
abstract
Asynchronous federated learning, which enables local clients to send their model update asynchronously to the server without waiting for others, has recently emerged for its improved efficiency and scalability over traditional synchronized federated learning. In this paper, we study how the asynchronous delay affects the convergence of asynchronous federated learning under non-i.i.d. distributed data across clients. Through the theoretical convergence analysis of one representative asynchronous federated learning algorithm under standard nonconvex stochastic settings, we show that the asynchronous delay can largely slow down the convergence, especially with high data heterogeneity. To further improve the convergence of asynchronous federated learning under heterogeneous data distributions, we propose a novel asynchronous federated learning method with a cached update calibration. Specifically, we let the server cache the latest update for each client and reuse these variables for calibrating the global update at each round. We theoretically prove the convergence acceleration for our proposed method under nonconvex stochastic settings. Extensive experiments on several vision and language tasks demonstrate our superior performances compared to other asynchronous federated learning baselines.
Yuanpu Cao, Jingcheng Wu
ICLR2
2024 Stealthy and Persistent Unalignment on Large Language Models via Backdoor Injections
abstract
Yuanpu Cao, Bochuan Cao, Jinghui Chen. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Yuanpu Cao, Bochuan Cao
NAACL-HLT1
2024 Personalized Steering of Large Language Models: Versatile Steering Vectors Through Bi-directional Preference Optimization
abstract
Researchers have been studying approaches to steer the behavior of Large Language Models (LLMs) and build personalized LLMs tailored for various applications. While fine-tuning seems to be a direct solution, it requires substantial computational resources and may significantly affect the utility of the original LLM. Recent endeavors have introduced more lightweight strategies, focusing on extracting ``steering vectors'' to guide the model's output toward desired behaviors by adjusting activations within specific layers of the LLM's transformer architecture. However, such steering vectors are directly extracted from the activations of human preference data and thus often lead to suboptimal results and occasional failures, especially in alignment-related scenarios. In this work, we propose an innovative approach that could produce more effective steering vectors through bi-directional preference optimization. Our method is designed to allow steering vectors to directly influence the generation probability of contrastive human preference data pairs, thereby offering a more precise representation of the target behavior. By carefully adjusting the direction and magnitude of the steering vector, we enabled personalized control over the desired behavior across a spectrum of intensities. Extensive experimentation across various open-ended generation tasks, particularly focusing on steering AI personas, has validated the efficacy of our approach. Moreover, we comprehensively investigate critical alignment-concerning scenarios, such as managing truthfulness, mitigating hallucination, and addressing jailbreaking attacks alongside their respective defenses. Remarkably, our method can still demonstrate outstanding steering effectiveness across these scenarios. Furthermore, we showcase the transferability of our steering vectors across different models/LoRAs and highlight the synergistic benefits of applying multiple vectors simultaneously. These findings significantly broaden the practicality and versatility of our proposed method.
Yuanpu Cao, Tianrong Zhang, Bochuan Cao, Ziyi Yin 0003, Lu Lin 0001, Fenglong Ma
NeurIPS1
2021 CTF: Anomaly Detection in High-Dimensional Time Series with Coarse-to-Fine Model Transfer
abstract
Anomaly detection is indispensable in modern IT infrastructure management. However, the dimension explosion problem of the monitoring data (large-scale machines, many key performance indicators, and frequent monitoring queries) causes a scalability issue to the existing algorithms. We propose a coarse-to-fine model transfer based framework CTF to achieve a scalable and accurate data-center-scale anomaly detection. CTF pre-trains a coarse-grained model, uses the model to extract and compress per-machine features to a distribution, clusters machines according to the distribution, and conducts model transfer to fine-tune per-cluster models for high accuracy. The framework takes advantage of clustering on the per-machine latent representation distribution, reusing the pre-trained model, and partial-layer model fine-tuning to boost the whole training efficiency. We also justify design choices such as the clustering algorithm and distance algorithm to achieve the best accuracy. We prototype CTF and experiment on production data to show its scalability and accuracy. We also release a labeling tool for multivariate time series and a labeled dataset to the research community.
Ya Su, Shenglin Zhang, Yuanpu Cao, Dan Pei, Wenfei Wu, Yongsu Zhang, Junliang Tang
INFOCOM4
2020 RLCard: A Platform for Reinforcement Learning in Card Games
abstract
We present RLCard, a Python platform for reinforcement learning research and development in card games. RLCard supports various card environments and several baseline algorithms with unified easy-to-use interfaces, aiming at bridging reinforcement learning and imperfect information games. The platform provides flexible configurations of state representation, action encoding, and reward design. RLCard also supports visualizations for algorithm debugging. In this demo, we showcase two representative environments and their visualization results. We conclude this demo with challenges and research opportunities brought by RLCard. A video is available on YouTube.
Daochen Zha, Kwei-Herng Lai, Songyi Huang, Yuanpu Cao, Keerthana Reddy, Juan Vargas, Alex Nguyen 0004, Ruzhe Wei, Xia Ben Hu
IJCAI4
2019 CoFlux: robustly correlating KPIs by fluctuations for service troubleshooting
abstract
Internet-based service companies monitor a large number of KPIs (Key Performance Indicators) to ensure their service quality and reliability. Correlating KPIs by fluctuations reveals interactions between KPIs under anomalous situations and can be extremely useful for service troubleshooting. However, such a KPI flux-correlation has been little studied so far in the domain of Internet service operations management. A major challenge is how to automatically and accurately separate fluctuations from normal variations in KPIs with different structural characteristics (such as seasonal, trend and stationary) for a large number of KPIs. In this paper, we propose CoFlux, an unsupervised approach, to automatically (without manual selection of algorithm fitting and parameter tuning) determine whether two KPIs are correlated by fluctuations, in what temporal order they fluctuate, and whether they fluctuate in the same direction. CoFlux's robust feature engineering and robust correlation score computation enable it to work well against the diverse KPI characteristics. Our extensive experiments have demonstrated that CoFlux achieves the best F1-Scores of 0.84 (0.90), 0.92 (0.95), 0.95 (0.99), in answering these three questions, in the two real datasets from a top global Internet company, respectively. Moreover, we showed that CoFlux is effective in assisting service troubleshooting through the applications of alert compression, recommending Top N causes, and constructing fluctuation propagation chains.
Ya Su, Youjian Zhao, Wentao Xia, Jiahao Bu, Jing Zhu 0007, Yuanpu Cao, Chenhao Niu, Yiyin Zhang, Zhaogang Wang, Dan Pei
IWQoS7