Zhenlan Ji

dblp:272/8251 · DBLP profile ↗
← Back
15ranked-venue papers
4as first author
14since 2021 · last 2026
0000-0003-3167-0480ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 8 · 4 first-author · 7 since 2021Security and privacy · 5 · 5 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Sok: Evaluating Jailbreak Guardrails for Large Language Models
abstract
Large Language Models (LLMs) have achieved remarkable progress, but their deployment has exposed critical vulnerabilities, particularly to jailbreak attacks that circumvent safety alignments. Guardrails--external defense mechanisms that monitor and control LLM interactions--have emerged as a promising solution. However, the current landscape of LLM guardrails is fragmented, lacking a unified taxonomy and comprehensive evaluation framework. In this Systematization of Knowledge (SoK) paper, we present the first holistic analysis of jailbreak guardrails for LLMs. We propose a novel, multi-dimensional taxonomy that categorizes guardrails along six key dimensions, and introduce a Security-Efficiency-Utility evaluation framework to assess their practical effectiveness. Through extensive analysis and experiments, we identify the strengths and limitations of existing guardrail approaches, provide insights into optimizing their defense mechanisms, and explore their universality across attack types. Our work offers a structured foundation for future research and development, aiming to guide the principled advancement and deployment of robust LLM guardrails. The code is available at https://github.com/xunguangwang/SoK4JailbreakGuardrails.
Xunguang Wang, Zhenlan Ji, Wenxuan Wang 0001, Zongjie Li, Daoyuan Wu, Shuai Wang 0011
SP2
2026 Instructta: instruction-tuned targeted attack for large vision-language models
abstract
Abstract Large vision-language models (LVLMs) have demonstrated their incredible capability in image understanding and response generation. However, this rich visual interaction also makes LVLMs vulnerable to adversarial examples. In this paper, we formulate a novel and practical targeted attack scenario that the adversary can only know the vision encoder of the victim LVLM, without the knowledge of its prompts (which are often proprietary for service providers and not publicly available) and its underlying large language model (LLM). This practical setting poses challenges to the cross-prompt and cross-model transferability of targeted adversarial attack, which aims to confuse the LVLM to output a response that is semantically similar to the attacker’s chosen target text. To this end, we propose an instruction-tuned targeted attack (dubbed I nstruct TA) to deliver the targeted adversarial attack on LVLMs with high transferability. Initially, we utilize a public text-to-image generative model to “reverse” the target response into a target image, and employ GPT-4 to infer a reasonable instruction $$\varvec{p}^\prime$$ p ′ from the target response. We then form a local surrogate model (sharing the same vision encoder with the victim LVLM) to extract instruction-aware features of an adversarial image example and the target image, and minimize the distance between these two features to optimize the adversarial example. To further improve the transferability with instruction tuning, we augment the instruction $$\varvec{p}^\prime$$ p ′ with instructions paraphrased from GPT-4. Extensive experiments on six victim LVLMs demonstrate the superiority of our proposed method in targeted attack performance and transferability. In particular, I nstruct TA achieves an attack success rate of 51.9% on BLIP-2, outperforming the strongest baseline by 10.5%, and consistently yields the highest attack success rates across all evaluated models.
Xunguang Wang, Pingchuan Ma 0004, Zhenlan Ji, Zongjie Li, Shuai Wang 0011, Weixi Gu
Cybersecur.3
2026 Reeq: Testing and Mitigating Ethically Inconsistent Suggestions of Large Language Models with Reflective Equilibrium
abstract
LLMs increasingly serve as general-purpose AI assistants in daily life, and their subtly unethical suggestions become a serious and real concern. It is demanding to test and mitigate such unethical suggestions from LLMs. Despite existing efforts to detect violations of “testable” facets of ethics (e.g., fairness testing), it is challenging to encode the full scope of ethics (e.g., justice, deontology) into a test oracle without human annotations or intervention. In this article, we take inspiration from reflective equilibrium, a modern moral reasoning method in moral and political philosophy, to guide our approach. Instead of seeking unethical suggestions in LLMs, we aim to identify behavioral inconsistency in LLMs’ ethics-related suggestions. These inconsistencies are anticipated to serve as a useful proxy and hint at unethical suggestions. We formulate reflective equilibrium in the form of fixed-point iteration, instantiate it as a novel test oracle, and also employ it to form a mitigation scheme for LLMs’ behavioral inconsistency on ethics-related inputs. To facilitate testing, we also create a comprehensive test suite, EthicsSuite , with 20K moral situations. In our study, we evaluate eight widely used LLMs. Our experiments reveal that LLMs are prone to ethical inconsistencies, with 81.22% of our test cases prompting ethically inconsistent suggestions on average. Our human evaluation suggests that the majority of these inconsistencies indeed manifest unethical biases. Our mitigation scheme effectively refines a significant number (80.1%) of these suggestions for commercial LLMs such as GPT-4 and Claude.
Pingchuan Ma 0004, Zhaoyu Wang 0006, Zongjie Li, Zhenlan Ji, Juergen Rahmel, Shuai Wang 0011
ACM Trans. Softw. Eng. Methodol.4
2025 The Phantom Menace in Crypto-Based PET-Hardened Deep Learning Models: Invisible Configuration-Induced Attacks
abstract
The increasing use of deep learning (DL) models has given rise to significant privacy concerns regarding training and inference data. To address these concerns, the community has increasingly adopted crypto-based privacy-enhancing technologies (CPET) like homomorphic encryption (HE), secure multi-party computation (MPC), and zero-knowledge proofs (ZKP). The integration of CPET with DL, often referred to as CPET-DL, is commonly facilitated by specialized frameworks like CrypTen, TenSEAL, and EZKL. These frameworks offer configurable parameters to balance model accuracy and computational efficiency during privacy-preserving operations. However, these configurations, while seemingly harmless, can introduce subtle vulnerabilities. The stealthy attacks induced by misconfigurations are hard to detect because 1) the plaintext models remain vulnerability-free, and 2) existing auditing tools are hardly applicable to CPET-hardened models. This creates a paradox: tools intended to protect privacy can be undermined through configuration manipulation.
Yiteng Peng, Dongwei Xiao, Zhibo Liu 0001, Zhenlan Ji, Daoyuan Wu, Shuai Wang 0011, Juergen Rahmel
CCS4
2025 Testing and Understanding Deviation Behaviors in FHE-Hardened Machine Learning Models
abstract
Fully homomorphic encryption (FHE) is a promising cryptographic primitive that enables secure computation over encrypted data. A primary use of FHE is to support privacypreserving machine learning (ML) on public cloud infrastructures. Despite the rapid development of FHE-based ML (or HE-ML), the community lacks a systematic understanding of their robustness. In this paper, we aim to systematically test and understand the deviation behaviors of HE-ML models, where the same input causes deviant outputs between FHE-hardened models and their plaintext versions, leading to completely incorrect model predictions. To effectively uncover deviation-triggering inputs under the constraints of expensive FHE computations, we design a novel differential testing tool called HEDIFF, which leverages the margin metric on the plaintext model as guidance to drive targeted testing on FHE models. For the identified deviation inputs, we further analyze them to determine whether they exhibit general noise patterns that are transferable. We evaluate HEDIFF using three popular HE-ML frameworks, covering 12 different combinations of models and datasets. HEDIFF successfully detected hundreds of deviation inputs across almost every tested FHE framework and model. We also quantitatively show that the identified deviation inputs are (visually) meaningful in comparison to regular inputs. Further schematic analysis reveals the root cause of these deviant inputs and allows us to generalize their noise patterns for more directed testing. Our work sheds light on enabling robust HE-ML for real-world usage.
Yiteng Peng, Daoyuan Wu, Zhibo Liu 0001, Dongwei Xiao, Zhenlan Ji, Juergen Rahmel, Shuai Wang 0011
ICSE5
2025 SelfDefend: LLMs Can Defend Themselves against Jailbreaking in a Practical Manner
Xunguang Wang, Daoyuan Wu, Zhenlan Ji, Zongjie Li, Pingchuan Ma 0004, Shuai Wang 0011, Yingjiu Li, Yang Liu 0003, Juergen Rahmel
USENIX Security Symposium3
2025 Guardrail: Automated Integrity Constraint Synthesis From Noisy Data
abstract
Data quality issues have been a long-standing challenge in the database community. Erroneous data can lead to incorrect query results, which in turn affect the credibility of the data-driven decisions. To circumvent this issue, a common practice is to discovery integrity constraints and enforce them on the data to ensure its quality. For instance, one can use constraints entailed by functional dependencies (FDs) to detect violations in the data. However, existing approaches fail to effectively discover them from noisy data. In this paper, we present a novel form of integrity constraints as a program under a domain-specific language (DSL) that can be used to detect and rectify errors in the data. On top of DSL, we propose an efficient synthesis algorithm that leverages the statistical structural properties of the data to generate the sketch of the program that significantly reduces the search space and speedup the synthesis process. To demonstrate the usefulness of our approach, we evaluate it on 12 real-world datasets for error detection. Then, we show that the synthesized integrity constraints can be used to solidify ML-integrated SQL queries over 48 queries, leading to an average reduction of 87% in the error rates. Our open-source artifact, including the G uardrail framework and the datasets, is available for the community to use [2].
Pingchuan Ma 0004, Zhaoyu Wang 0006, Zhenlan Ji, Zongjie Li, Shuai Wang 0011
Proc. ACM Manag. Data3
2025 Privacy-preserving and Verifiable Causal Prescriptive Analytics
abstract
Prescriptive analytics seeks to identify optimal interventions for achieving desired outcomes, with causal inference playing a pivotal role in assessing intervention impacts on complex systems. However, existing approaches frequently neglect critical data privacy considerations and provide no means to verify the integrity of their recommendations. These limitations hinder its adoption in high-stakes domains such as healthcare and finance. In this paper, we introduce, zkCLEAR, a zero-knowledge proof (ZKP)-based C ausal Inference ( LEA rning and R easoning) framework for privacy-preserving and verifiable prescriptive analytics. Our solution allows data owners or service providers to cryptographically prove the validity of prescriptive conclusions derived from causal analysis without disclosing sensitive source data or proprietary causal models. We develop a suite of ZKP-friendly causal operators to build efficient causal modules, including structure learning, parameter learning, probabilistic inference, and counterfactual reasoning. To optimize performance, we also introduce a workflow decomposition strategy to facilitate efficient proof generation for complex workloads. We demonstrate the utility of zkCLEAR through three real-world applications. The framework faithfully follows the behavior of non-ZKP counterparts, with moderate overheads for privacy and verifiability. Additionally, we evaluate its efficiency and scalability using real-world datasets. It shows up to a 35.1× speedup in proof generation time and a 214.5× reduction in proof size compared to current general-purpose ZKP systems.
Zhaoyu Wang 0006, Pingchuan Ma 0004, Zhantong Xue, Yanbo Dai, Zhenlan Ji, Shuai Wang 0011
Proc. ACM Manag. Data5
2024 Enabling Runtime Verification of Causal Discovery Algorithms with Automated Conditional Independence Reasoning
abstract
Causal discovery is a powerful technique for identifying causal relationships among variables in data. It has been widely used in various applications in software engineering. Causal discovery extensively involves conditional independence (CI) tests. Hence, its output quality highly depends on the performance of CI tests, which can often be unreliable in practice. Moreover, privacy concerns arise when excessive CI tests are performed.
Pingchuan Ma 0004, Zhenlan Ji, Peisen Yao, Shuai Wang 0011, Kui Ren 0001
ICSE2
2023 CC: Causality-Aware Coverage Criterion for Deep Neural Networks
abstract
Deep neural network (DNN) testing approaches have grown fast in recent years to test the correctness and robustness of DNNs. In particular, DNN coverage criteria are frequently used to evaluate the quality of a test suite, and a number of coverage criteria based on neuron-wise, layer-wise, and path-/trace-wise coverage patterns have been published to date. However, we see that existing criteria are insufficient to represent how one neuron would influence subsequent neurons; hence, we lack a concept of how neurons, when functioning as causes and effects, might jointly make a DNN prediction. Given recent advances in interpreting DNN internals using causal inference, we present the first causality-aware DNN coverage criterion, which evaluates a test suite by quantifying the extent to which the suite provides new causal relations for testing DNNs. Performing standard causal inference on DNNs presents both theoretical and practical hurdles. We introduce CC (causal coverage), a practical and efficient coverage criterion that integrates a set of optimizations using DNN domain-specific knowledge. We illustrate the efficacy of CC using diverse, real-world inputs and adversarial inputs, such as adversarial examples (AEs) and backdoor inputs. We demonstrate that CC outperforms previous DNN criteria under various settings with moderate cost.
Zhenlan Ji, Pingchuan Ma 0004, Yuanyuan Yuan 0001, Shuai Wang 0011
ICSE1
2023 Perfce: Performance Debugging on Databases with Chaos Engineering-Enhanced Causality Analysis
abstract
Debugging performance anomalies in databases is challenging. Causal inference techniques enable qualitative and quantitative root cause analysis of performance downgrades. Nevertheless, causality analysis is challenging in practice, particularly due to limited observability. Recently, chaos engineering (CE) has been applied to test complex software systems. CE frameworks mutate chaos variables to inject catastrophic events (e.g., network slowdowns) to stress-test these software systems. The systems under chaos stress are then tested (e.g., via differential testing) to check if they retain normal functionality, such as returning correct SQL query outputs even under stress. To date, CE is mainly employed to aid software testing. This paper identifies the novel usage of CE in diagnosing performance anomalies in databases. Our framework, PERFCE, has two phases - offline and online. The offline phase learns statistical models of a database using both passive observations and proactive chaos experiments. The online phase diagnoses the root cause of performance anomalies from both qualitative and quantitative aspects on-the-fly. In evaluation, Perfce outperformed previous works on synthetic datasets and is highly accurate and moderately expensive when analyzing real-world (distributed) databases like MySQL and TiDB.
Zhenlan Ji, Pingchuan Ma 0004, Shuai Wang 0011
ASE1
2023 Causality-Aided Trade-Off Analysis for Machine Learning Fairness
abstract
There has been an increasing interest in enhancing the fairness of machine learning (ML). Despite the growing number of fairness-improving methods, we lack a systematic understanding of the trade-offs among factors considered in the ML pipeline when fairness-improving methods are applied. This understanding is essential for developers to make informed decisions regarding the provision of fair ML services. Nonetheless, it is extremely difficult to analyze the trade-offs when there are multiple fairness parameters and other crucial metrics involved, coupled, and even in conflict with one another. This paper uses causality analysis as a principled method for analyzing trade-offs between fairness parameters and other crucial metrics in ML pipelines. To practically and effectively conduct causality analysis, we propose a set of domain-specific optimizations to facilitate accurate causal discovery and a unified, novel interface for trade-off analysis based on well-established causal inference methods. We conduct a comprehensive empirical study using three real-world datasets on a collection of widely-used fairness-improving techniques. Our study obtains actionable suggestions for users and developers of fair ML. We further demonstrate the versatile usage of our approach in selecting the optimal fairness-improving method, paving the way for more ethical and socially responsible AI technologies.
Zhenlan Ji, Pingchuan Ma 0004, Shuai Wang 0011, Yanhui Li 0001
ASE1
2022 Unlearnable Examples: Protecting Open-Source Software from Unauthorized Neural Code Learning
abstract
The vast volume of "free" code maintained on open-source code management systems significantly simplifies the process of producing and sharing open-source software.Recently, we have seen a growing trend in which these open-source software is being used for neural code learning without authorization.Note that open-source software does not necessarily imply "unrestricted usage," e.g., software under the BSD license requires users to retain the copyright notice and credit the software's developers.The unauthorized use of software for (commercial) neural code learning models has raised copyright concerns.This paper, for the first time, provides approaches for protecting opensource software from unauthorized neural code learning via unlearnable examples.Our proposed technique applies a set of lightweight transformations toward a program before it is open-source released.When these transformed programs are used to train models, they mislead the model into learning the unnecessary knowledge of programs, then fail the model to complete original programs.The transformation methods are sophisticatedly designed to ensure that they do not impair the general readability of protected programs, nor do they entail a huge cost.We focus on code autocompletion as a representative downstream task of unauthorized neural code learning.We demonstrate highly encouraging and cost-effective protection against neural code autocompletion.
Zhenlan Ji, Pingchuan Ma 0004, Shuai Wang 0011
SEKE1
2022 NoLeaks: Differentially Private Causal Discovery Under Functional Causal Model
abstract
Causal inference is widely used in clinical research, economic analysis, and other fields. As is the case with many statistical data, the findings of causal discovery (i.e., causal graph) might leak demographic information of participants. For example, a causal link between one genome and a rare disease can reveal the participation of a minority patient in genome-ide association studies. To date, differential privacy has served as the de facto foundation for guaranteeing the privacy of causal discovery algorithms. However, existing approaches to protecting causal discovery from privacy leakage rely heavily on private conditional independence tests, which generate a considerable amount of noise and are thus prone to inaccuracy. As a result of their limited accuracy and scalability, they are insufficient for non-trivial datasets (e.g., those with more than ten variables). In this paper, we advocate a novel focus on enforcing privacy for causal discovery algorithms based on functional causal models. First, we propose NOLEAKS, a differentially private causal discovery algorithm, which manifests both high accuracy and efficiency compared with prior works. Second, we design a quasi-Newton numerical optimization algorithm for solving NOLEAKS in a highly efficient way. Third, we evaluate NOLEAKS using both public benchmarks and synthetic data. We observe that NOLEAKS achieves comparable performance or even surpasses the state-of-the-art (non-private) approaches. We also find encouraging results that NOLEAKS can smoothly scale to large datasets, on which existing works would fail. Through a case study and a downstream application, we observe encouraging results on the versatile usages of NOLEAKS.
Pingchuan Ma 0004, Zhenlan Ji, Qi Pang, Shuai Wang 0011
IEEE Trans. Inf. Forensics Secur.2
2020 An Empirical Study on Issue Knowledge Transfer from Python to R for Machine Learning Software
Wen-Chin Huang, Zhenlan Ji
SEKE2