Jiahao Yu 0001

dblp:238/6241-1 · DBLP profile ↗
← Back
13ranked-venue papers
5as first author
12since 2021 · last 2026
0009-0007-4919-0967ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Security and privacy · 5 · 3 first-author · 5 since 2021Computer networks · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 PromptFuzz: Harnessing Fuzzing Techniques for Robust Testing of Prompt Injection in LLMs
abstract
Large Language Models (LLMs) have gained widespread use in various applications due to their powerful capability to generate human-like text. However, prompt injection attacks, which involve overwriting a model’s original instructions with malicious prompts to manipulate the generated text, have raised significant concerns about the security and reliability of LLMs. In this paper, we propose PromptFuzz, a novel testing framework that leverages fuzzing techniques to systematically assess the robustness of LLMs against prompt injection attacks. Inspired by software fuzzing, PromptFuzz selects promising seed prompts and generates a diverse set of prompt injections to evaluate the target LLM’s resilience. PromptFuzz operates in two stages: theprepare phase, which involves selecting promising initial seeds and collecting few-shot examples, and thefocus phase, which uses the collected examples to generate diverse, high-quality prompt injections. By deploying the generated attack prompts from PromptFuzz in a real-world competition, we achieved the 7th ranking out of over 4000 participants (top 0.14%) within 2 hours, demonstrating PromptFuzz’s effectiveness compared to experienced human attackers. Additionally, we also deploy the generated attack prompts on 50 popular LLM-integrated online applications, including those from Coze and OpenAI, and found that 92% of them can be exploited by PromptFuzz. We also run PromptFuzz on 15 online LLM-based resume judging applications and found that 13 of these applications’ responses can be hijacked by PromptFuzz.
Yangguang Shao, Jiahao Yu 0001, Hanwen Miao, Gaopeng Gou, Zhen Li 0011, Junzheng Shi
IEEE Trans. Inf. Forensics Secur.2
2025 GPO: Learning from Critical Steps to Improve LLM Reasoning
abstract
Large language models (LLMs) are increasingly used in various domains, showing impressive potential on various tasks. Recently, reasoning LLMs have been proposed to improve the \textit{reasoning} or \textit{thinking} capabilities of LLMs to solve complex problems. Despite the promising results of reasoning LLMs, enhancing the multi-step reasoning capabilities of LLMs still remains a significant challenge. While existing optimization methods have advanced the LLM reasoning capabilities, they often treat reasoning trajectories as a whole, without considering the underlying critical steps within the trajectory. In this paper, we introduce \textbf{G}uided \textbf{P}ivotal \textbf{O}ptimization (GPO), a novel fine-tuning strategy that dives into the reasoning process to enable more effective improvements. GPO first identifies the `critical step' within a reasoning trajectory - a point that the model must carefully proceed so as to succeed at the problem. We locate the critical step by estimating the advantage function. GPO then resets the policy to the critical step and samples the new rollout and prioritizes learning process on those rollouts. This focus allows the model to learn more effectively from pivotal moments within the reasoning process to improve the reasoning performance. We demonstrate that GPO is not a standalone method, but rather a general strategy that can be integrated with various optimization methods to improve reasoning performance. Besides theoretical analysis, our experiments across challenging reasoning benchmarks show that GPO can consistently and significantly enhances the performance of existing optimization methods, showcasing its effectiveness and generalizability in improving LLM reasoning by concentrating on pivotal moments within the generation process.
Jiahao Yu 0001, Zelei Cheng, Xian Wu 0007, Xinyu Xing 0001
NeurIPS1
2025 BlockScan: Detecting Anomalies in Blockchain Transactions
abstract
We propose BlockScan, a customized Transformer for anomaly detection in blockchain transactions. Unlike existing methods that rely on rule-based systems or directly apply off-the-shelf large language models (LLMs), BlockScan introduces a series of customized designs to effectively model the unique data structure of blockchain transactions. First, a blockchain transaction is multi-modal, containing blockchain-specific tokens, texts, and numbers. We design a novel modularized tokenizer to handle these multi-modal inputs, balancing the information across different modalities. Second, we design a customized masked language modeling mechanism for pretraining the Transformer architecture, incorporating RoPE embedding and FlashAttention for handling longer sequences. Finally, we design a novel anomaly detection method based on the model outputs. We further provide theoretical analysis for the detection method of our system. Extensive evaluations on Ethereum and Solana transactions demonstrate BlockScan's exceptional capability in anomaly detection while maintaining a low false positive rate. Remarkably, BlockScan is the only method that successfully detects anomalous transactions on Solana with high accuracy, whereas all other approaches achieved very low or zero detection recall scores. This work sets a new benchmark for applying Transformer-based approaches in blockchain data analysis.
Jiahao Yu 0001, Xian Wu 0007, Wenbo Guo 0002, Xinyu Xing 0001
NeurIPS1
2025 Mind the Inconspicuous: Revealing the Hidden Weakness in Aligned LLMs' Refusal Boundaries
Jiahao Yu 0001, Haozheng Luo, Jerry Yao-Chieh Hu, Yan Chen 0004, Wenbo Guo 0002, Han Liu 0001, Xinyu Xing 0001
USENIX Security Symposium1
2025 PATCHAGENT: A Practical Program Repair Agent Mimicking Human Expertise
Zheng Yu 0003, Yuhang Wu 0003, Jiahao Yu 0001, Meng Xu 0025, Dongliang Mu, Yan Chen 0004, Xinyu Xing 0001
USENIX Security Symposium4
2024 RICE: Breaking Through the Training Bottlenecks of Reinforcement Learning with Explanation
abstract
Deep reinforcement learning (DRL) is playing an increasingly important role in real-world applications. However, obtaining an optimally performing DRL agent for complex tasks, especially with sparse rewards, remains a significant challenge. The training of a DRL agent can be often trapped in a bottleneck without further progress. In this paper, we propose RICE, an innovative refining scheme for reinforcement learning that incorporates explanation methods to break through the training bottlenecks. The high-level idea of RICE is to construct a new initial state distribution that combines both the default initial states and critical states identified through explanation methods, thereby encouraging the agent to explore from the mixed initial states. Through careful design, we can theoretically guarantee that our refining scheme has a tighter sub-optimality bound. We evaluate RICE in various popular RL environments and real-world applications. The results demonstrate that RICE significantly outperforms existing refining schemes in enhancing agent performance.
Zelei Cheng, Xian Wu 0007, Jiahao Yu 0001, Sabrina Yang, Gang Wang 0011, Xinyu Xing 0001
ICML3
2024 Soft-Label Integration for Robust Toxicity Classification
abstract
Toxicity classification in textual content remains a significant problem. Data with labels from a single annotator fall short of capturing the diversity of human perspectives. Therefore, there is a growing need to incorporate crowdsourced annotations for training an effective toxicity classifier. Additionally, the standard approach to training a classifier using empirical risk minimization (ERM) may fail to address the potential shifts between the training set and testing set due to exploiting spurious correlations. This work introduces a novel bi-level optimization framework that integrates crowdsourced annotations with the soft-labeling technique and optimizes the soft-label weights by Group Distributionally Robust Optimization (GroupDRO) to enhance the robustness against out-of-distribution (OOD) risk. We theoretically prove the convergence of our bi-level optimization algorithm. Experimental results demonstrate that our approach outperforms existing baseline methods in terms of both average and worst-group accuracy, confirming its effectiveness in leveraging crowdsourced annotations to achieve more effective and robust toxicity classification.
Zelei Cheng, Xian Wu 0007, Jiahao Yu 0001, Xin-Qiang Cai, Xinyu Xing 0001
NeurIPS3
2024 LLM-Fuzzer: Scaling Assessment of Large Language Model Jailbreaks
Jiahao Yu 0001, Xingwei Lin, Zheng Yu 0003, Xinyu Xing 0001
USENIX Security Symposium1
2023 StateMask: Explaining Deep Reinforcement Learning through State Mask
abstract
Despite the promising performance of deep reinforcement learning (DRL) agents in many challenging scenarios, the black-box nature of these agents greatly limits their applications in critical domains. Prior research has proposed several explanation techniques to understand the deep learning-based policies in RL. Most existing methods explain why an agent takes individual actions rather than pinpointing the critical steps to its final reward. To fill this gap, we propose StateMask, a novel method to identify the states most critical to the agent's final reward. The high-level idea of StateMask is to learn a mask net that blinds a target agent and forces it to take random actions at some steps without compromising the agent's performance. Through careful design, we can theoretically ensure that the masked agent performs similarly to the original agent. We evaluate StateMask in various popular RL environments and show its superiority over existing explainers in explanation fidelity. We also show that StateMask has better utilities, such as launching adversarial attacks and patching policy errors.
Zelei Cheng, Xian Wu 0007, Jiahao Yu 0001, Wenhai Sun, Wenbo Guo 0002, Xinyu Xing 0001
NeurIPS3
2023 AIRS: Explanation for Deep Reinforcement Learning based Security Applications
Jiahao Yu 0001, Wenbo Guo 0002, Gang Wang 0011, Ting Wang 0006, Xinyu Xing 0001
USENIX Security Symposium1
2023 Matrix Gaussian Mechanisms for Differentially-Private Learning
abstract
The wide deployment of machine learning algorithms has become a severe threat to user data privacy. As the learning data is of high dimensionality and high orders, preserving its privacy is intrinsically hard. Conventional differential privacy mechanisms often incur significant utility decline as they are designed for scalar values from the start. We recognize that it is because conventional approaches do not take the data structural information into account, and fail to provide sufficient privacy or utility. As the main novelty of this work, we proposeMatrix Gaussian Mechanism(MGM), a new$ (\epsilon,\delta)$-differential privacy mechanism for preserving learning data privacy. By imposing the unimodal distributions on the noise, we introduce two mechanisms based on MGM with an improved utility. We further show that with the utility space available, the proposed mechanisms can be instantiated with optimized utility, and has a closed-form solution scalable to large-scale problems. We experimentally show that our mechanisms, applied to privacy-preserving federated learning, are superior than the state-of-the-art differential privacy mechanisms in utility.
Jungang Yang 0002, Liyao Xiang, Jiahao Yu 0001, Xinbing Wang, Bin Guo 0001, Zhetao Li, Baochun Li
IEEE Trans. Mob. Comput.3
2021 Speedup Robust Graph Structure Learning with Low-Rank Information
abstract
Recent studies have shown that graph neural networks (GNNs) are vulnerable to unnoticeable adversarial perturbations, which largely confines their deployment in many safety-critical domains. Robust graph structure learning has been proposed to improve the GNN performance in the face of adversarial attacks. In particular, the low-rank methods are utilized to purify the perturbed graphs. However, these methods are mostly computationally expensive with O(n3) time complexity and O(n2) space complexity. We propose LRGNN, a fast and robust graph structure learning framework, which exploits the low-rank property as prior knowledge to speed up optimization. To eliminate adversarial perturbation, LRGNN decouples the adjacency matrix into a low-rank component and a sparse one, and learns by minimizing the rank of the first part while suppressing the second part. Its sparse variant is formed to reduce the memory footprint further. Experimental results on various attack settings have shown LRGNN acquires comparable robustness with the state-of-the-art much more efficiently, boasting a significant advantage on large-scale graphs.
Hui Xu 0011, Liyao Xiang, Jiahao Yu 0001, Xinbing Wang
CIKM3
2020 Voiceprint Mimicry Attack Towards Speaker Verification System in Smart Home
abstract
The advancement of voice controllable systems (VC-Ses) has dramatically affected our daily lifestyle and catalyzed the smart home's deployment. Currently, most VCSes exploit automatic speaker verification (ASV) to prevent various voice attacks (e.g., replay attack). In this study, we present VMask, a novel and practical voiceprint mimicry attack that could fool ASV in smart home and inject the malicious voice command disguised as a legitimate user. The key observation behind VMask is that the deep learning models utilized by ASV are vulnerable to the subtle perturbations in the voice input space. To generate these subtle perturbations, VMask leverages the idea of adversarial examples. Then by adding the subtle perturbations to the recordings from an arbitrary speaker, VMask can mislead the ASV into classifying the crafted speech samples, which mirror the former speaker for human, as the targeted victim. Moreover, psychoacoustic masking is employed to manipulate the adversarial perturbations under human perception threshold, thus making victim unaware of ongoing attacks. We validate the effectiveness of VMask by performing comprehensive experiments on both grey box (VGGVox) and black box (Microsoft Azure Speaker Verification) ASVs. Additionally, a real-world case study on Apple HomeKit proves the VMask's practicability on smart home platforms.
Yan Meng 0001, Jiahao Yu 0001, Chong Xiang 0001, Brandon Falk, Haojin Zhu
INFOCOM3