VLDB 2026 Research / reviewers in the wild / expert
Kaijie Zhu
dblp:56/7058
· DBLP profile ↗
23ranked-venue papers
8as first author
20since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 5 first-author · 13 since 2021Databases, data management, data science and information retrieval · 5 · 3 first-author · 3 since 2021Computer networks · 3 · 2 since 2021Security and privacy · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Understanding and Mitigating Bias Inheritance in LLM-based Data Augmentation on Downstream TasksabstractMiaomiao Li, Hao Chen, Yang Wang, Tingyuan Zhu, Weijia Zhang, Kaijie Zhu, Kam-Fai Wong, Jindong Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Hao Chen 0102, Tingyuan Zhu, Kaijie Zhu, Kam-Fai Wong, Jindong Wang 0001 |
ACL (1) | 6 |
| 2026 | STAIR: Towards structure-aware inference of autonomous system relationships using graph neural networks
Zekun Tao, Kaijie Zhu, Peng Zhang 0011, Yue Chen 0018 |
Comput. Networks | 3 |
| 2026 | On topology and time: efficient evaluation for temporal-clique subgraph queriesabstractAbstract We investigate temporal-clique subgraph pattern matching, where edges must both form a specific topological sub-structure and temporally overlap within a specified window. This problem has widespread applications across domains including social networks, life sciences, smart cities, and telecommunications. However, existing subgraph matching techniques are inefficient at processing such queries that combine both temporal and structural constraints. We propose a novel approach that effectively leverages both topological and temporal selectivities of the query to significantly improve processing performance. Our solution introduces key innovations across the query processing pipeline, including a specialized multi-way join operator, an optimized query planner, and an accurate cardinality estimator. Through additional optimizations, we further enhance the efficiency of our approach. Extensive experiments demonstrate that our method substantially outperforms state-of-the-art techniques while requiring minimal additional storage overhead. Kaijie Zhu, Shichang Ding, George Fletcher 0001, Nikolay Yakovets |
VLDB J. | 1 |
| 2025 | MELON: Provable Defense Against Indirect Prompt Injection Attacks in AI AgentsabstractRecent research has explored that LLM agents are vulnerable to indirect prompt injection (IPI) attacks, where malicious tasks embedded in tool-retrieved information can redirect the agent to take unauthorized actions. Existing defenses against IPI have significant limitations: either require essential model training resources, lack effectiveness against sophisticated attacks, or harm the normal utilities. We present MELON (Masked re-Execution and TooL comparisON), a novel IPI defense. Our approach builds on the observation that under a successful attack, the agent’s next action becomes less dependent on user tasks and more on malicious tasks. Following this, we design MELON to detect attacks by re-executing the agent’s trajectory with a masked user prompt modified through a masking function. We identify an attack if the actions generated in the original and masked executions are similar. We also include three key designs to reduce the potential false positives and false negatives. Extensive evaluation on the IPI benchmark AgentDojo demonstrates that MELON outperforms SOTA defenses in both attack prevention and utility preservation. Moreover, we show that combining MELON with a SOTA prompt augmentation defense (denoted as MELON-Aug) further improves its performance. We also conduct a detailed ablation study to validate our key designs. Code is available at https://github.com/kaijiezhu11/MELON. Kaijie Zhu, Xianjun Yang, Jindong Wang 0001, Wenbo Guo 0002, William Yang Wang |
ICML | 1 |
| 2025 | Co-PatcheR: Collaborative Software Patching with Component-specific Small Reasoning ModelsabstractMotivated by the success of general‑purpose large language models (LLMs) in software patching, recent works started to train specialized patching models. Most works trained one model to handle the end‑to‑end patching pipeline (including issue localization, patch generation, and patch validation). However, it is hard for a small model to handle all tasks, as different sub-tasks have different workflows and require different expertise. As such, by using a 70 billion model, SOTA methods can only reach up to 41% resolved rate on SWE-bench-Verified. Motivated by the collaborative nature, we propose Co-PatcheR, the first collaborative patching system with small and specialized reasoning models for individual components. Our key technique novelties are the specific task designs and training recipes. First, we train a model for localization and patch generation. Our localization pinpoints the suspicious lines through a two-step procedure, and our generation combines patch generation and critique. We then propose a hybrid patch validation that includes two models for crafting issue-reproducing test cases with and without assertions and judging patch correctness, followed by a majority vote-based patch selection. Through extensive evaluation, we show that Co-PatcheR achieves 46% resolved rate on SWE-bench-Verified with only 3 x 14B models. This makes Co-PatcheR the best patcher with specialized models, requiring the least training resources and the smallest models. We conduct a comprehensive ablation study to validate our recipes, as well as our choice of training data number, model size, and testing-phase scaling strategy. Yuheng Tang, Hongwei Li 0025, Kaijie Zhu, Michael Yang, Yangruibo Ding, Wenbo Guo 0002 |
NeurIPS | 3 |
| 2025 | Personalized Safety in LLMs: A Benchmark and A Planning-Based Agent ApproachabstractLarge language models (LLMs) typically generate identical or similar responses for all users given the same prompt, posing serious safety risks in high-stakes applications where user vulnerabilities differ widely.
Existing safety evaluations primarily rely on context-independent metrics—such as factuality, bias, or toxicity—overlooking the fact that the same response may carry divergent risks depending on the user's background or condition.
We introduce ``personalized safety'' to fill this gap and present PENGUIN—a benchmark comprising 14,000 scenarios across seven sensitive domains with both context-rich and context-free variants. Evaluating six leading LLMs, we demonstrate that personalized user information significantly improves safety scores by 43.2%, confirming the effectiveness of personalization in safety alignment. However, not all context attributes contribute equally to safety enhancement. To address this, we develop RAISE—a training-free, two-stage agent framework that strategically acquires user-specific background. RAISE improves safety scores by up to 31.6% over six vanilla LLMs, while maintaining a low interaction cost of just 2.7 user queries on average. Our findings highlight the importance of selective information gathering in safety-critical domains and offer a practical solution for personalizing LLM responses without model retraining. This work establishes a foundation for safety research that adapts to individual user contexts rather than assuming a universal harm standard. Edward Sun, Kaijie Zhu, Jianxun Lian, José Hernández-Orallo, Aylin Caliskan, Jindong Wang 0001 |
NeurIPS | 3 |
| 2025 | AECR: Automatic attack technique intelligence extraction based on fine-tuned large language model
Minghao Chen 0003, Kaijie Zhu, Qingjun Yuan, Yuefei Zhu |
Comput. Secur. | 2 |
| 2024 | AgentReview: Exploring Peer Review Dynamics with LLM AgentsabstractPeer review is fundamental to the integrity and advancement of scientific publication.Traditional methods of peer review analyses often rely on exploration and statistics of existing peer review data, which do not adequately address the multivariate nature of the process, account for the latent variables, and are further constrained by privacy concerns due to the sensitive nature of the data.We introduce AGENTREVIEW, the first large language model (LLM) based peer review simulation framework, which effectively disentangles the impacts of multiple latent factors and addresses the privacy issue.Our study reveals significant insights, including a notable 37.1% variation in paper decisions due to reviewers' biases, supported by sociological theories such as the social influence theory, altruism fatigue, and authority bias.We believe that this study could offer valuable insights to improve the design of peer review mechanisms.Our code is available at https://github.com/Ahren09/AgentReview. Yiqiao Jin, Qinlin Zhao, Hao Chen 0102, Kaijie Zhu, Yijia Xiao, Jindong Wang 0001 |
EMNLP | 5 |
| 2024 | An IP Anti-geolocation Method Based on Constructed Landmarks
Enshang Lu, Shichang Ding, Chunfang Yang, Daofu Gong, Kaijie Zhu, Xiangyang Luo 0001 |
ICDF2C (2) | 6 |
| 2024 | DyVal: Dynamic Evaluation of Large Language Models for Reasoning TasksabstractLarge language models (LLMs) have achieved remarkable performance in various evaluation benchmarks. However, concerns are raised about potential data contamination in their considerable volume of training corpus. Moreover, the static nature and fixed complexity of current benchmarks may inadequately gauge the advancing capabilities of LLMs.
In this paper, we introduce DyVal, a general and flexible protocol for dynamic evaluation of LLMs. Based on our framework, we build graph-informed DyVal by leveraging the structural advantage of directed acyclic graphs to dynamically generate evaluation samples with controllable complexities. DyVal generates challenging evaluation sets on reasoning tasks including mathematics, logical reasoning, and algorithm problems. We evaluate various LLMs ranging from Flan-T5-large to GPT-3.5-Turbo and GPT-4. Experiments show that LLMs perform worse in DyVal-generated evaluation samples with different complexities, highlighting the significance of dynamic evaluation.
We also analyze the failure cases and results of different prompting methods.
Moreover, DyVal-generated samples are not only evaluation sets, but also helpful data for fine-tuning to improve the performance of LLMs on existing benchmarks.
We hope that DyVal can shed light on future evaluation research of LLMs. Code is available at: https://github.com/microsoft/promptbench. Kaijie Zhu, Jiaao Chen, Jindong Wang 0001, Neil Zhenqiang Gong, Diyi Yang, Xing Xie 0001 |
ICLR | 1 |
| 2024 | The Good, The Bad, and Why: Unveiling Emotions in Generative AIabstractEmotion significantly impacts our daily behaviors and interactions. While recent generative AI models, such as large language models, have shown impressive performance in various tasks, it remains unclear whether they truly comprehend emotions and why. This paper aims to address this gap by incorporating psychological theories to gain a holistic understanding of emotions in generative AI models. Specifically, we propose three approaches: 1) EmotionPrompt to enhance AI model performance, 2) EmotionAttack to impair AI model performance, and 3) EmotionDecode to explain the effects of emotional stimuli, both benign and malignant. Through extensive experiments involving language and multi-modal models on semantic understanding, logical reasoning, and generation tasks, we demonstrate that both textual and visual EmotionPrompt can boost the performance of AI models while EmotionAttack can hinder it. More importantly, EmotionDecode reveals that AI models can comprehend emotional stimuli akin to the mechanism of dopamine in the human brain. Our work heralds a novel avenue for exploring psychology to enhance our understanding of generative AI models, thus boosting the research and development of human-AI collaboration and mitigating potential risks. Jindong Wang 0001, Yixuan Zhang 0001, Kaijie Zhu, Wenxin Hou, Jianxun Lian, Qiang Yang 0001, Xing Xie 0001 |
ICML | 4 |
| 2024 | CompeteAI: Understanding the Competition Dynamics of Large Language Model-based AgentsabstractLarge language models (LLMs) have been widely used as agents to complete different tasks, such as personal assistance or event planning. Although most of the work has focused on cooperation and collaboration between agents, little work explores competition, another important mechanism that promotes the development of society and economy. In this paper, we seek to examine the competition dynamics in LLM-based agents. We first propose a general framework for studying the competition between agents. Then, we implement a practical competitive environment using GPT-4 to simulate a virtual town with two types of agents, including restaurant agents and customer agents. Specifically, the restaurant agents compete with each other to attract more customers, where competition encourages them to transform, such as cultivating new operating strategies. Simulation experiments reveal several interesting findings at the micro and macro levels, which align well with existing market and sociological theories. We hope that the framework and environment can be a promising testbed to study the competition that fosters understanding of society. Code is available at: https://github.com/microsoft/competeai. Qinlin Zhao, Jindong Wang 0001, Yixuan Zhang 0001, Yiqiao Jin, Kaijie Zhu, Hao Chen 0102, Xing Xie 0001 |
ICML | 5 |
| 2024 | Dynamic Evaluation of Large Language Models by Meta Probing AgentsabstractEvaluation of large language models (LLMs) has raised great concerns in the community due to the issue of data contamination. Existing work designed evaluation protocols using well-defined algorithms for specific tasks, which cannot be easily extended to diverse scenarios. Moreover, current evaluation benchmarks can only provide the overall benchmark results and cannot support a fine-grained and multifaceted analysis of LLMs’ abilities. In this paper, we propose meta probing agents (MPA), a general dynamic evaluation protocol inspired by psychometrics to evaluate LLMs. MPA designs the probing and judging agents to automatically transform an original evaluation problem into a new one following psychometric theory on three basic cognitive abilities: language understanding, problem solving, and domain knowledge. These basic abilities are also dynamically configurable, allowing multifaceted analysis. We conducted extensive evaluations using MPA and found that most LLMs achieve poorer performance, indicating room for improvement. Our multifaceted analysis demonstrated the strong correlation between the basic abilities and an implicit Mattew effect on model size, i.e., larger models possess stronger correlations of the abilities. MPA can also be used as a data augmentation approach to enhance LLMs. Code is available at: https://github.com/microsoft/promptbench. Kaijie Zhu, Jindong Wang 0001, Qinlin Zhao, Ruochen Xu, Xing Xie 0001 |
ICML | 1 |
| 2024 | Flatter Minima of Loss Landscapes Correspond with Strong Corruption Robustness
Liqun Zhong, Kaijie Zhu, Ge Yang 0002 |
ICPR (1) | 2 |
| 2024 | PromptBench: A Unified Library for Evaluation of Large Language ModelsabstractThe evaluation of large language models (LLMs) is crucial to assess their performance and mitigate potential security risks. In this paper, we introduce PromptBench, a unified library to evaluate LLMs. It consists of several key components that can be easily used and extended by researchers: prompt construction, prompt engineering, dataset and model loading, adversarial prompt attack, dynamic evaluation protocols, and analysis tools. PromptBench is designed as an open, general, and flexible codebase for research purpose. It aims to facilitate original study in creating new benchmarks, deploying downstream applications, and designing new evaluation protocols. The code is available at: https://github.com/microsoft/promptbench and will be continuously supported. Kaijie Zhu, Qinlin Zhao, Hao Chen 0102, Jindong Wang 0001, Xing Xie 0001 |
J. Mach. Learn. Res. | 1 |
| 2024 | A Survey on Evaluation of Large Language ModelsabstractLarge language models (LLMs) are gaining increasing popularity in both academia and industry, owing to their unprecedented performance in various applications. As LLMs continue to play a vital role in both research and daily use, their evaluation becomes increasingly critical, not only at the task level, but also at the society level for better understanding of their potential risks. Over the past years, significant efforts have been made to examine LLMs from various perspectives. This paper presents a comprehensive review of these evaluation methods for LLMs, focusing on three key dimensions: what to evaluate , where to evaluate , and how to evaluate . Firstly, we provide an overview from the perspective of evaluation tasks, encompassing general natural language processing tasks, reasoning, medical usage, ethics, education, natural and social sciences, agent applications, and other areas. Secondly, we answer the ‘where’ and ‘how’ questions by diving into the evaluation methods and benchmarks, which serve as crucial components in assessing the performance of LLMs. Then, we summarize the success and failure cases of LLMs in different tasks. Finally, we shed light on several future challenges that lie ahead in LLMs evaluation. Our aim is to offer invaluable insights to researchers in the realm of LLMs evaluation, thereby aiding the development of more proficient LLMs. Our key point is that evaluation should be treated as an essential discipline to better assist the development of LLMs. We consistently maintain the related open-source materials at: https://github.com/MLGroupJLU/LLM-eval-survey Yupeng Chang, Jindong Wang 0001, Yuan Wu 0002, Linyi Yang, Kaijie Zhu, Hao Chen 0102, Xiaoyuan Yi, Cunxiang Wang, Yidong Wang 0003, Wei Ye 0004, Yue Zhang 0004, Yi Chang 0001, Philip S. Yu, Qiang Yang 0001, Xing Xie 0001 |
ACM Trans. Intell. Syst. Technol. | 6 |
| 2023 | Improving Generalization of Adversarial Training via Robust Critical Fine-TuningabstractDeep neural networks are susceptible to adversarial examples, posing a significant security risk in critical applications. Adversarial Training (AT) is a well-established technique to enhance adversarial robustness, but it often comes at the cost of decreased generalization ability. This paper proposes Robustness Critical Fine-Tuning (RiFT), a novel approach to enhance generalization without compromising adversarial robustness. The core idea of RiFT is to exploit the redundant capacity for robustness by fine-tuning the adversarially trained model on its non-robust-critical module. To do so, we introduce module robust criticality (MRC), a measure that evaluates the significance of a given module to model robustness under worst-case weight perturbations. Using this measure, we identify the module with the lowest MRC value as the non-robust-critical module and fine-tune its weights to obtain fine-tuned weights. Subsequently, we linearly interpolate between the adversarially trained weights and fine-tuned weights to derive the optimal fine-tuned model weights. We demonstrate the efficacy of RiFT on ResNet18, ResNet34, and WideResNet34-10 models trained on CIFAR10, CIFAR100, and Tiny-ImageNet datasets. Our experiments show that RiFT can significantly improve both generalization and out-of-distribution robustness by around 1.5% while maintaining or even slightly enhancing adversarial robustness. Code is available at https://github.com/Immortalise/RiFT. Kaijie Zhu, Xixu Hu, Jindong Wang 0001, Xing Xie 0001, Ge Yang 0002 |
ICCV | 1 |
| 2023 | CBFLNet: Cross-boundary feature learning for large-scale point cloud segmentation
Bingyao Wang, Chengyang Li 0001, Kaijie Zhu |
Eng. Appl. Artif. Intell. | 5 |
| 2022 | Anomaly Detection as a Service: An Outsourced Anomaly Detection Scheme for Blockchain in a Privacy-Preserving MannerabstractAttacks against blockchain networks have proliferated in recent years. Due to its immense economic value, Bitcoin has been subject to numerous malicious theft activities through the exchange platforms. This poses a severe threat to the credibility of the entire Bitcoin ecosystem. Therefore, it is necessary to provide detection and prediction services of malicious events for Bitcoin Exchanges to prevent them in a precise and timely manner. Meanwhile, preserving the privacy of transaction data to prevent de-anonymization attacks during the detection process is also of great importance. In this paper, we present a general framework for privacy-preserving anomaly detection in blockchain networks. Based on this framework, we propose ADaaS, an anomaly detection service scheme that adopts a supervised machine learning model and achieves privacy preservation by using vector homomorphic encryption and matrix perturbation strategies. We also analyze the security, communication and computation costs of ADaaS. Experimental results demonstrate that ADaaS can achieve high detection effectiveness while providing privacy guarantees and is applicable in real scenarios of detecting Bitcoin transactions due to its reasonable efficiency. Yuhan Song, Fushan Wei, Kaijie Zhu, Yuefei Zhu |
IEEE Trans. Netw. Serv. Manag. | 3 |
| 2021 | Leveraging Temporal and Topological Selectivities in Temporal-clique Subgraph Query ProcessingabstractWe study the problem of temporal-clique subgraph pattern matching. In such patterns, edges are required to jointly overlap in time within a given temporal window in addition to forming a topological sub-structure. This problem arises in many application domains, e.g., in social networks, life sciences, smart cities, telecommunications, and others. State-of-the-art subgraph matching techniques, however, are shown to be limited and inefficient in processing queries with both temporal and topological constraints. We propose an approach that takes full advantage of both topological and temporal selectivities during the processing of temporal-clique subgraph queries. Additionally, we investigate a number of optimizations that can be introduced into our approach to improve its efficiency. Our experimental results demonstrate that our approach outperforms the existing methods by a wide margin at a small additional storage cost. Kaijie Zhu, George Fletcher 0001, Nikolay Yakovets |
ICDE | 1 |
| 2019 | Secure Cryptography Infrastructures in the CloudabstractInformation systems are deployed in clouds as virtual machines (VMs) for better agility, elasticity and reliability. It is necessary to safekeep their cryptographic keys, e.g., the private keys used in TLS and SSH, against various attacks. However, existing virtualization solutions do not improve the cryptography facilities of in-cloud systems. This paper presents SECRIN, a secure cryptography infrastructure for VMs in the cloud. SECRIN is composed of a) virtual cryptographic devices implemented in VM monitors (VMMs), and b) a device management tool integrated in the virtualization management system. A virtual device receives requests from VMs, computes with cryptographic keys within the VMM and returns results. The keys appear only in the VMM's memory space, so that they are kept secret even if the VMs were compromised. With the management tool, the operator of virtualization management systems assigns virtual cryptographic devices to a VM as well as other resources, while the tenant (or owner) of a VM still holds proper controls on the keys. The virtual devices work compatibly with live migration, and the cryptographic computations are not interrupted when the VMs are moving from a host to another. We develop the SECRIN prototype with KVM-QEMU and oVirt. Experimental results show that, it works compatibly with existing virtualization solutions, provides reliable cryptographic computing services for applications, and is secure against attacks happening in VMs. Dawei Chu, Kaijie Zhu, Quanwei Cai 0001, Jingqiang Lin 0001, Fengjun Li, Le Guan, Lingchen Zhang |
GLOBECOM | 2 |
| 2019 | Scalable temporal clique enumerationabstractWe study the problem of enumeration of all k-sized subsets of temporal events that mutually overlap at some point in a query time window. This problem arises in many application domains, e.g., in social networks, life sciences, smart cities, telecommunications, and others. We propose a start time index (STI) approach that overcomes the efficiency bottlenecks of current methods which are based on 2-way join algorithms to enumerate temporal k-cliques. Additionally, we investigate how precomputed checkpoints can be used to further improve the efficiency of STI. Our experimental results demonstrate that STI outperforms the state of the art by a wide margin and that our checkpointing strategies are effective. Kaijie Zhu, George Fletcher 0001, Nikolay Yakovets, Odysseas Papapetrou, Yuqing Wu |
SSTD | 1 |
| 2019 | Cluster-preserving sampling from fully-dynamic streaming graphs
Kaijie Zhu, Yulong Pei, George Fletcher 0001, Mykola Pechenizkiy |
Inf. Sci. | 2 |